{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/16897"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/16897","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Mining informative sentences in multi-product documents with PLSA","abstract":"In this thesis, we study the problem of fetching informative sentences from multi-product documents. A multi-product document is defined as a document that mentions about multiple products, which is often for comparison purpose. An informative sentence is defined as a sentence that provides the characteristics of the product. we propose to use Probabilistic Latent Semantic Analysis (PLSA) to mine informative sentences given a single multi-product document. By applying PLSA in a multi-product document, it simultaneously solves three problems regarding to multi-product document: 1. Separate the document based on the products. 2. Fetch the informative key words for each product. 3. Fetch the informative sentences. The proposed method can mine the product information from a single multi-product document, which is quite different from previous works in opinion mining. Experiment results show that the high probability words in the word distribution of each product do discover the characteristics of the product. The results also reveal that the sentences which contain more these keywords are more informative. Practical applications of this method are: 1. Product comparison. 2. Facilitate user in reading multi-product document, which is to apply data mining in human computer interaction (HCI).","abstract_html":"In this thesis, we study the problem of fetching informative sentences from multi-product documents. A multi-product document is defined as a document that mentions about multiple products, which is often for comparison purpose. An informative sentence is defined as a sentence that provides the characteristics of the product. we propose to use Probabilistic Latent Semantic Analysis (PLSA) to mine informative sentences given a single multi-product document. By applying PLSA in a multi-product document, it simultaneously solves three problems regarding to multi-product document: 1. Separate the document based on the products. 2. Fetch the informative key words for each product. 3. Fetch the informative sentences. The proposed method can mine the product information from a single multi-product document, which is quite different from previous works in opinion mining. Experiment results show that the high probability words in the word distribution of each product do discover the characteristics of the product. The results also reveal that the sentences which contain more these keywords are more informative. Practical applications of this method are: 1. Product comparison. 2. Facilitate user in reading multi-product document, which is to apply data mining in human computer interaction (HCI).","abstract_has_math":false,"creators":["Hwang, Hsiang-Yeh"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Zhai, ChengXiang"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2010,"date_issued":"2010-08-20T18:01:10Z","date_published":"2010-08-20T18:01:10Z","updated_at":"2026-07-22T22:25:09Z","subjects":["multi-product document","single-product document","informative sentences","informative words (ie. keywords)","product name entity"],"languages":["en"],"rights":["Copyright 2010 Hsiang-Yeh Hwang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/16897","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhai, ChengXiang"]},{"key":"dc:creator","label":"Author","values":["Hwang, Hsiang-Yeh"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2010-08-20T18:01:10Z","2010-08"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["multi-product document","single-product document","informative sentences","informative words (ie. keywords)","product name entity"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2010 Hsiang-Yeh Hwang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/16897"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["In this thesis, we study the problem of fetching informative sentences from multi-product documents. A multi-product document is defined as a document that mentions about multiple products, which is often for comparison purpose. An informative sentence is defined as a sentence that provides the characteristics of the product. we propose to use Probabilistic Latent Semantic Analysis (PLSA) to mine informative sentences given a single multi-product document. By applying PLSA in a multi-product document, it simultaneously solves three problems regarding to multi-product document: 1. Separate the document based on the products. 2. Fetch the informative key words for each product. 3. Fetch the informative sentences. The proposed method can mine the product information from a single multi-product document, which is quite different from previous works in opinion mining. Experiment results show that the high probability words in the word distribution of each product do discover the characteristics of the product. The results also reveal that the sentences which contain more these keywords are more informative. Practical applications of this method are: 1. Product comparison. 2. Facilitate user in reading multi-product document, which is to apply data mining in human computer interaction (HCI).","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2010-07-20T13:33:41Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Hsiang-Yeh_Hwang.pdf: 2735829 bytes, checksum: 6596d844cc5ec2fe1384bcfd7965db80 (MD5)","Made available in DSpace on 2010-08-20T18:01:10Z (GMT). No. of bitstreams: 3 Hsiang-Yeh_Hwang.pdf: 2735829 bytes, checksum: 6596d844cc5ec2fe1384bcfd7965db80 (MD5) Hwang_Hsiang-Yeh.pdf: 2735824 bytes, checksum: 1b3935938aaf6a38430880d6832ef66d (MD5) license.txt: 4061 bytes, checksum: a793f64dbfe95c45162b22c5887d7e25 (MD5)"]},{"key":"dc:title","label":"Title","values":["Mining informative sentences in multi-product documents with PLSA"]}]}],"canonical_facts":{"dc:contributor":["Zhai, ChengXiang"],"dc:creator":["Hwang, Hsiang-Yeh"],"dc:date":["2010-08-20T18:01:10Z","2010-08"],"dc:description":["In this thesis, we study the problem of fetching informative sentences from multi-product documents. A multi-product document is defined as a document that mentions about multiple products, which is often for comparison purpose. An informative sentence is defined as a sentence that provides the characteristics of the product. we propose to use Probabilistic Latent Semantic Analysis (PLSA) to mine informative sentences given a single multi-product document. By applying PLSA in a multi-product document, it simultaneously solves three problems regarding to multi-product document: 1. Separate the document based on the products. 2. Fetch the informative key words for each product. 3. Fetch the informative sentences. The proposed method can mine the product information from a single multi-product document, which is quite different from previous works in opinion mining. Experiment results show that the high probability words in the word distribution of each product do discover the characteristics of the product. The results also reveal that the sentences which contain more these keywords are more informative. Practical applications of this method are: 1. Product comparison. 2. Facilitate user in reading multi-product document, which is to apply data mining in human computer interaction (HCI).","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2010-07-20T13:33:41Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Hsiang-Yeh_Hwang.pdf: 2735829 bytes, checksum: 6596d844cc5ec2fe1384bcfd7965db80 (MD5)","Made available in DSpace on 2010-08-20T18:01:10Z (GMT). No. of bitstreams: 3 Hsiang-Yeh_Hwang.pdf: 2735829 bytes, checksum: 6596d844cc5ec2fe1384bcfd7965db80 (MD5) Hwang_Hsiang-Yeh.pdf: 2735824 bytes, checksum: 1b3935938aaf6a38430880d6832ef66d (MD5) license.txt: 4061 bytes, checksum: a793f64dbfe95c45162b22c5887d7e25 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/16897"],"dc:language":["en"],"dc:rights":["Copyright 2010 Hsiang-Yeh Hwang"],"dc:subject":["multi-product document","single-product document","informative sentences","informative words (ie. keywords)","product name entity"],"dc:title":["Mining informative sentences in multi-product documents with PLSA"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:09Z"}