{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/16875"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/16875","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Object search: supporting structured queries in web search engines","abstract":"As the web evolves, increasing quantities of structured information is embedded in web pages in disparate formats. For example, a digital camera’s description may include its price and megapixels whereas a professor’s description may include her name, university, and research interests. Both types of pages may include additional ambiguous information. General search engines (GSEs) do not support queries over these types of data because they ignore the web document semantics. Conversely, describing requisite semantics through structured queries into databases populated by information extraction (IE) techniques are expensive and not easily adaptable to new domains. This paper describes a methodology for rapidly developing search engines capable of answering structured queries over unstructured corpora by utilizing machine learning to avoid explicit IE. We empirically show that with minimum additional human effort, our system outperforms a GSE with respect to structured queries with clear object semantics.","abstract_html":"As the web evolves, increasing quantities of structured information is embedded in web pages in disparate formats. For example, a digital camera’s description may include its price and megapixels whereas a professor’s description may include her name, university, and research interests. Both types of pages may include additional ambiguous information. General search engines (GSEs) do not support queries over these types of data because they ignore the web document semantics. Conversely, describing requisite semantics through structured queries into databases populated by information extraction (IE) techniques are expensive and not easily adaptable to new domains. This paper describes a methodology for rapidly developing search engines capable of answering structured queries over unstructured corpora by utilizing machine learning to avoid explicit IE. We empirically show that with minimum additional human effort, our system outperforms a GSE with respect to structured queries with clear object semantics.","abstract_has_math":false,"creators":["Pham, Cuong K."],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Chang, Kevin C-C."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2010,"date_issued":"2010-08-20T18:00:32Z","date_published":"2010-08-20T18:00:32Z","updated_at":"2026-07-22T22:25:09Z","subjects":["semantics","web search engines","learning to rank","information retrieval","structured query","object search"],"languages":["en"],"rights":["Copyright 2010 by Cuong Kim Pham. All rights reserved."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/16875","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Chang, Kevin C-C."]},{"key":"dc:creator","label":"Author","values":["Pham, Cuong K."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2010-08-20T18:00:32Z","2010-08"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["semantics","web search engines","learning to rank","information retrieval","structured query","object search"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2010 by Cuong Kim Pham. All rights reserved."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/16875"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["As the web evolves, increasing quantities of structured information is embedded in web pages in disparate formats. For example, a digital camera’s description may include its price and megapixels whereas a professor’s description may include her name, university, and research interests. Both types of pages may include additional ambiguous information. General search engines (GSEs) do not support queries over these types of data because they ignore the web document semantics. Conversely, describing requisite semantics through structured queries into databases populated by information extraction (IE) techniques are expensive and not easily adaptable to new domains. This paper describes a methodology for rapidly developing search engines capable of answering structured queries over unstructured corpora by utilizing machine learning to avoid explicit IE. We empirically show that with minimum additional human effort, our system outperforms a GSE with respect to structured queries with clear object semantics.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2010-07-16T19:03:54Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 3 Pham_KimCuong.pdf: 560512 bytes, checksum: fc9825c9a5562e70b1c3a29975b2ac67 (MD5) msthesis-objectsearch-kimpham-jul10.pdf: 560512 bytes, checksum: fc9825c9a5562e70b1c3a29975b2ac67 (MD5) Pham_KimCuong.pdf: 567002 bytes, checksum: 7161e8f3649a226cff168671046bc4e3 (MD5)","Made available in DSpace on 2010-08-20T18:00:32Z (GMT). No. of bitstreams: 4 Pham_KimCuong.pdf: 560512 bytes, checksum: fc9825c9a5562e70b1c3a29975b2ac67 (MD5) msthesis-objectsearch-kimpham-jul10.pdf: 560512 bytes, checksum: fc9825c9a5562e70b1c3a29975b2ac67 (MD5) 1_Pham_KimCuong.pdf: 567002 bytes, checksum: 7161e8f3649a226cff168671046bc4e3 (MD5) license.txt: 4060 bytes, checksum: d15f4bb274332f93701bcbc67b940bf5 (MD5)"]},{"key":"dc:title","label":"Title","values":["Object search: supporting structured queries in web search engines"]}]}],"canonical_facts":{"dc:contributor":["Chang, Kevin C-C."],"dc:creator":["Pham, Cuong K."],"dc:date":["2010-08-20T18:00:32Z","2010-08"],"dc:description":["As the web evolves, increasing quantities of structured information is embedded in web pages in disparate formats. For example, a digital camera’s description may include its price and megapixels whereas a professor’s description may include her name, university, and research interests. Both types of pages may include additional ambiguous information. General search engines (GSEs) do not support queries over these types of data because they ignore the web document semantics. Conversely, describing requisite semantics through structured queries into databases populated by information extraction (IE) techniques are expensive and not easily adaptable to new domains. This paper describes a methodology for rapidly developing search engines capable of answering structured queries over unstructured corpora by utilizing machine learning to avoid explicit IE. We empirically show that with minimum additional human effort, our system outperforms a GSE with respect to structured queries with clear object semantics.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2010-07-16T19:03:54Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 3 Pham_KimCuong.pdf: 560512 bytes, checksum: fc9825c9a5562e70b1c3a29975b2ac67 (MD5) msthesis-objectsearch-kimpham-jul10.pdf: 560512 bytes, checksum: fc9825c9a5562e70b1c3a29975b2ac67 (MD5) Pham_KimCuong.pdf: 567002 bytes, checksum: 7161e8f3649a226cff168671046bc4e3 (MD5)","Made available in DSpace on 2010-08-20T18:00:32Z (GMT). No. of bitstreams: 4 Pham_KimCuong.pdf: 560512 bytes, checksum: fc9825c9a5562e70b1c3a29975b2ac67 (MD5) msthesis-objectsearch-kimpham-jul10.pdf: 560512 bytes, checksum: fc9825c9a5562e70b1c3a29975b2ac67 (MD5) 1_Pham_KimCuong.pdf: 567002 bytes, checksum: 7161e8f3649a226cff168671046bc4e3 (MD5) license.txt: 4060 bytes, checksum: d15f4bb274332f93701bcbc67b940bf5 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/16875"],"dc:language":["en"],"dc:rights":["Copyright 2010 by Cuong Kim Pham. All rights reserved."],"dc:subject":["semantics","web search engines","learning to rank","information retrieval","structured query","object search"],"dc:title":["Object search: supporting structured queries in web search engines"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:09Z"}