{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/34506"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/34506","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Utilizing multiple entities from collection of unstructured documents in constructing attribute-value pairs","abstract":"Attribute-value pairs, or NVP is defined as extracting words expressing characteristics of entity and associating the said words with word or phrases that best describe the attributes. Applications for NVP arise in various related area such as sentiment analysis, populating and checking for errors in relational database to a broader text information area such as QA systems, search and review modeling. We propose an unsupervised method to identify the properties of entities represented as NVP from unstructured documents. Other approaches that extract NVP usually uti- lize supervised or semi-supervised approaches on structured or semi-structured documents. Benefits of such approaches lie in that they tend to have higher accuracy than unsuper- vised approaches on unstructured documents. Furthermore, supervised approaches are more suited to distinguishing attribute words to that of value words than unsupervised approaches on unstructured documents. The biggest drawback with the said methods however, is that training data may not always be available and not all documents can be thought of as being unstructured. We first proposes in this thesis an approach to extracting and distinguishing attribute words and value words from unstructured documents. Since entities of the same class share similar attributes, we propose that the identification of relevant attributes should be done across entities belonging to the same class, and demonstrate that this can lead to a significant performance gain in attribute extraction, even when only documents describing a modest number of entities per class is available. We then propose a way to evaluate the accuracy of attribute-value pairs automatically, allowing for quantitative comparison between different systems that is more consistent and cost-effective than manual evaluations. These were used in evaluating summarization or comparing ontologies. However, these techniques have not been utilized in evaluating NVP. Both the automated and manual evaluations show that our system outperforms a comparison system.","abstract_html":"Attribute-value pairs, or NVP is defined as extracting words expressing characteristics of entity and associating the said words with word or phrases that best describe the attributes. Applications for NVP arise in various related area such as sentiment analysis, populating and checking for errors in relational database to a broader text information area such as QA systems, search and review modeling. We propose an unsupervised method to identify the properties of entities represented as NVP from unstructured documents. Other approaches that extract NVP usually uti- lize supervised or semi-supervised approaches on structured or semi-structured documents. Benefits of such approaches lie in that they tend to have higher accuracy than unsuper- vised approaches on unstructured documents. Furthermore, supervised approaches are more suited to distinguishing attribute words to that of value words than unsupervised approaches on unstructured documents. The biggest drawback with the said methods however, is that training data may not always be available and not all documents can be thought of as being unstructured. We first proposes in this thesis an approach to extracting and distinguishing attribute words and value words from unstructured documents. Since entities of the same class share similar attributes, we propose that the identification of relevant attributes should be done across entities belonging to the same class, and demonstrate that this can lead to a significant performance gain in attribute extraction, even when only documents describing a modest number of entities per class is available. We then propose a way to evaluate the accuracy of attribute-value pairs automatically, allowing for quantitative comparison between different systems that is more consistent and cost-effective than manual evaluations. These were used in evaluating summarization or comparing ontologies. However, these techniques have not been utilized in evaluating NVP. Both the automated and manual evaluations show that our system outperforms a comparison system.","abstract_has_math":false,"creators":["Cho, Hyun Duk"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Zhai, ChengXiang"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2012,"date_issued":"2012-09-18T21:20:36Z","date_published":"2012-09-18T21:20:36Z","updated_at":"2026-07-22T22:25:31Z","subjects":["attribute extraction","(attribute-value pair) nvp","value extraction","evaluation"],"languages":["en"],"rights":["Copyright 2012 Hyun Duk Cho"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/34506","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhai, ChengXiang"]},{"key":"dc:creator","label":"Author","values":["Cho, Hyun Duk"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2012-09-18T21:20:36Z","2014-09-18T10:00:36Z","2012-08"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["attribute extraction","(attribute-value pair) nvp","value extraction","evaluation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2012 Hyun Duk Cho"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/34506"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Attribute-value pairs, or NVP is defined as extracting words expressing characteristics of entity and associating the said words with word or phrases that best describe the attributes. Applications for NVP arise in various related area such as sentiment analysis, populating and checking for errors in relational database to a broader text information area such as QA systems, search and review modeling. We propose an unsupervised method to identify the properties of entities represented as NVP from unstructured documents. Other approaches that extract NVP usually uti- lize supervised or semi-supervised approaches on structured or semi-structured documents. Benefits of such approaches lie in that they tend to have higher accuracy than unsuper- vised approaches on unstructured documents. Furthermore, supervised approaches are more suited to distinguishing attribute words to that of value words than unsupervised approaches on unstructured documents. The biggest drawback with the said methods however, is that training data may not always be available and not all documents can be thought of as being unstructured. We first proposes in this thesis an approach to extracting and distinguishing attribute words and value words from unstructured documents. Since entities of the same class share similar attributes, we propose that the identification of relevant attributes should be done across entities belonging to the same class, and demonstrate that this can lead to a significant performance gain in attribute extraction, even when only documents describing a modest number of entities per class is available. We then propose a way to evaluate the accuracy of attribute-value pairs automatically, allowing for quantitative comparison between different systems that is more consistent and cost-effective than manual evaluations. These were used in evaluating summarization or comparing ontologies. However, these techniques have not been utilized in evaluating NVP. Both the automated and manual evaluations show that our system outperforms a comparison system.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2012-07-18T12:33:11Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Cho_Hyun Duk.pdf: 2072774 bytes, checksum: ea58d070d1356785a47752fd0743c850 (MD5) Cho_Hyun Duk.pdf: 2072783 bytes, checksum: 31d2ff3868d2bfdae74a5587a0cd6048 (MD5)","Made available in DSpace on 2012-09-18T21:20:36Z (GMT). No. of bitstreams: 2 Cho_Hyun Duk.pdf: 2072781 bytes, checksum: 495aeb7ac01b697b26b88dbdfaf39f49 (MD5) license.txt: 4056 bytes, checksum: 823b28fc057fca872a25cd2b62e092d7 (MD5)","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by Seth Robbins (srobbins@illinois.edu) on 2012-09-18T21:21:21Z Item is restricted until 2014-09-18T21:21:01Z","Restriction data tranferred 2014-07-01T11:35:13-05:00 Original Data Group with Access UIUC Users [automated] Release Date: 2014-09-18 16:21:01 UTC Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 34780 on 2014-09-18T10:00:36Z."]},{"key":"dc:title","label":"Title","values":["Utilizing multiple entities from collection of unstructured documents in constructing attribute-value pairs"]}]}],"canonical_facts":{"dc:contributor":["Zhai, ChengXiang"],"dc:creator":["Cho, Hyun Duk"],"dc:date":["2012-09-18T21:20:36Z","2014-09-18T10:00:36Z","2012-08"],"dc:description":["Attribute-value pairs, or NVP is defined as extracting words expressing characteristics of entity and associating the said words with word or phrases that best describe the attributes. Applications for NVP arise in various related area such as sentiment analysis, populating and checking for errors in relational database to a broader text information area such as QA systems, search and review modeling. We propose an unsupervised method to identify the properties of entities represented as NVP from unstructured documents. Other approaches that extract NVP usually uti- lize supervised or semi-supervised approaches on structured or semi-structured documents. Benefits of such approaches lie in that they tend to have higher accuracy than unsuper- vised approaches on unstructured documents. Furthermore, supervised approaches are more suited to distinguishing attribute words to that of value words than unsupervised approaches on unstructured documents. The biggest drawback with the said methods however, is that training data may not always be available and not all documents can be thought of as being unstructured. We first proposes in this thesis an approach to extracting and distinguishing attribute words and value words from unstructured documents. Since entities of the same class share similar attributes, we propose that the identification of relevant attributes should be done across entities belonging to the same class, and demonstrate that this can lead to a significant performance gain in attribute extraction, even when only documents describing a modest number of entities per class is available. We then propose a way to evaluate the accuracy of attribute-value pairs automatically, allowing for quantitative comparison between different systems that is more consistent and cost-effective than manual evaluations. These were used in evaluating summarization or comparing ontologies. However, these techniques have not been utilized in evaluating NVP. Both the automated and manual evaluations show that our system outperforms a comparison system.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2012-07-18T12:33:11Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Cho_Hyun Duk.pdf: 2072774 bytes, checksum: ea58d070d1356785a47752fd0743c850 (MD5) Cho_Hyun Duk.pdf: 2072783 bytes, checksum: 31d2ff3868d2bfdae74a5587a0cd6048 (MD5)","Made available in DSpace on 2012-09-18T21:20:36Z (GMT). No. of bitstreams: 2 Cho_Hyun Duk.pdf: 2072781 bytes, checksum: 495aeb7ac01b697b26b88dbdfaf39f49 (MD5) license.txt: 4056 bytes, checksum: 823b28fc057fca872a25cd2b62e092d7 (MD5)","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by Seth Robbins (srobbins@illinois.edu) on 2012-09-18T21:21:21Z Item is restricted until 2014-09-18T21:21:01Z","Restriction data tranferred 2014-07-01T11:35:13-05:00 Original Data Group with Access UIUC Users [automated] Release Date: 2014-09-18 16:21:01 UTC Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 34780 on 2014-09-18T10:00:36Z."],"dc:identifier":["http://hdl.handle.net/2142/34506"],"dc:language":["en"],"dc:rights":["Copyright 2012 Hyun Duk Cho"],"dc:subject":["attribute extraction","(attribute-value pair) nvp","value extraction","evaluation"],"dc:title":["Utilizing multiple entities from collection of unstructured documents in constructing attribute-value pairs"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:31Z"}