{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/81771"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/81771","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Large Scale Information Integration on the Web: Finding, Understanding and Querying Web Databases","abstract":"\"The Web has been rapidly \"\"deepened\"\" by myriad searchable databases online, where data are hidden behind query interfaces. Guarding data behind there, such query interfaces are the \"\"entrances\"\" or \"\"doors\"\" to the deep Web. To open this door to the deep Web, we have been building the MetaQuerier system---for both exploring (to find) and integrating (to query) databases on the Web through their query interfaces. To find Web databases, we need to provide search functionalities that dynamically discover databases relevant to user's information needs. To query those Web databases, we need to \"\"understand\"\" what a query interface says---i.e., what query capabilities a source supports through its interface, in terms of specifiable conditions. Further, to help users query \"\"alternative\"\" sources, we need to mediate heterogeneous query capabilities across different sources discovered on-the-fly. Finally, to process queries submitted to a database, we need to design efficient query processing techniques. To address those challenges, this thesis presents several key components in MetaQuerier system: First, a search facility searches for useful databases by their schemas; Second, form extractor extracts query capabilities of databases by applying a best-effort parsing approach based on hidden syntax; Third, form assistant translates queries across pairs of interfaces on-the-fly by deploying a light-weight, domain-based translation framework. Fourth, OPT* framework processes ranked queries by a k constraint optimization problem. We evaluate our techniques upon real databases on the Web. The experiment results show the promise of our system.\"","abstract_html":"&quot;The Web has been rapidly &quot;&quot;deepened&quot;&quot; by myriad searchable databases online, where data are hidden behind query interfaces. Guarding data behind there, such query interfaces are the &quot;&quot;entrances&quot;&quot; or &quot;&quot;doors&quot;&quot; to the deep Web. To open this door to the deep Web, we have been building the MetaQuerier system---for both exploring (to find) and integrating (to query) databases on the Web through their query interfaces. To find Web databases, we need to provide search functionalities that dynamically discover databases relevant to user&#x27;s information needs. To query those Web databases, we need to &quot;&quot;understand&quot;&quot; what a query interface says---i.e., what query capabilities a source supports through its interface, in terms of specifiable conditions. Further, to help users query &quot;&quot;alternative&quot;&quot; sources, we need to mediate heterogeneous query capabilities across different sources discovered on-the-fly. Finally, to process queries submitted to a database, we need to design efficient query processing techniques. To address those challenges, this thesis presents several key components in MetaQuerier system: First, a search facility searches for useful databases by their schemas; Second, form extractor extracts query capabilities of databases by applying a best-effort parsing approach based on hidden syntax; Third, form assistant translates queries across pairs of interfaces on-the-fly by deploying a light-weight, domain-based translation framework. Fourth, OPT* framework processes ranked queries by a k constraint optimization problem. We evaluate our techniques upon real databases on the Web. The experiment results show the promise of our system.&quot;","abstract_has_math":false,"creators":["Zhang, Zhen"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Chang, Kevin C."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-09-25T20:20:23Z","date_published":"2015-09-25T20:20:23Z","updated_at":"2026-07-22T22:26:16Z","subjects":["Computer Science"],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["(MiAaPQ)AAI3270067"],"render_values":[{"text":"(MiAaPQ)AAI3270067","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/2142/81771","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Chang, Kevin C."]},{"key":"dc:creator","label":"Author","values":["Zhang, Zhen"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2015-09-25T20:20:23Z","10000-01-01","2007"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/81771","(MiAaPQ)AAI3270067"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["\"The Web has been rapidly \"\"deepened\"\" by myriad searchable databases online, where data are hidden behind query interfaces. Guarding data behind there, such query interfaces are the \"\"entrances\"\" or \"\"doors\"\" to the deep Web. To open this door to the deep Web, we have been building the MetaQuerier system---for both exploring (to find) and integrating (to query) databases on the Web through their query interfaces. To find Web databases, we need to provide search functionalities that dynamically discover databases relevant to user's information needs. To query those Web databases, we need to \"\"understand\"\" what a query interface says---i.e., what query capabilities a source supports through its interface, in terms of specifiable conditions. Further, to help users query \"\"alternative\"\" sources, we need to mediate heterogeneous query capabilities across different sources discovered on-the-fly. Finally, to process queries submitted to a database, we need to design efficient query processing techniques. To address those challenges, this thesis presents several key components in MetaQuerier system: First, a search facility searches for useful databases by their schemas; Second, form extractor extracts query capabilities of databases by applying a best-effort parsing approach based on hidden syntax; Third, form assistant translates queries across pairs of interfaces on-the-fly by deploying a light-weight, domain-based translation framework. Fourth, OPT* framework processes ranked queries by a k constraint optimization problem. We evaluate our techniques upon real databases on the Web. The experiment results show the promise of our system.\"","Made available in DSpace on 2015-09-25T20:20:23Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 3270067.pdf: 4592455 bytes, checksum: 1177db785b9c79a6df5e99c3fb06da92 (MD5) Previous issue date: 2007","Embargo set by: Seth Robbins for item 83052 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","148 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2007."]},{"key":"dc:title","label":"Title","values":["Large Scale Information Integration on the Web: Finding, Understanding and Querying Web Databases"]}]}],"canonical_facts":{"dc:contributor":["Chang, Kevin C."],"dc:creator":["Zhang, Zhen"],"dc:date":["2015-09-25T20:20:23Z","10000-01-01","2007"],"dc:description":["\"The Web has been rapidly \"\"deepened\"\" by myriad searchable databases online, where data are hidden behind query interfaces. Guarding data behind there, such query interfaces are the \"\"entrances\"\" or \"\"doors\"\" to the deep Web. To open this door to the deep Web, we have been building the MetaQuerier system---for both exploring (to find) and integrating (to query) databases on the Web through their query interfaces. To find Web databases, we need to provide search functionalities that dynamically discover databases relevant to user's information needs. To query those Web databases, we need to \"\"understand\"\" what a query interface says---i.e., what query capabilities a source supports through its interface, in terms of specifiable conditions. Further, to help users query \"\"alternative\"\" sources, we need to mediate heterogeneous query capabilities across different sources discovered on-the-fly. Finally, to process queries submitted to a database, we need to design efficient query processing techniques. To address those challenges, this thesis presents several key components in MetaQuerier system: First, a search facility searches for useful databases by their schemas; Second, form extractor extracts query capabilities of databases by applying a best-effort parsing approach based on hidden syntax; Third, form assistant translates queries across pairs of interfaces on-the-fly by deploying a light-weight, domain-based translation framework. Fourth, OPT* framework processes ranked queries by a k constraint optimization problem. We evaluate our techniques upon real databases on the Web. The experiment results show the promise of our system.\"","Made available in DSpace on 2015-09-25T20:20:23Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 3270067.pdf: 4592455 bytes, checksum: 1177db785b9c79a6df5e99c3fb06da92 (MD5) Previous issue date: 2007","Embargo set by: Seth Robbins for item 83052 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","148 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2007."],"dc:identifier":["http://hdl.handle.net/2142/81771","(MiAaPQ)AAI3270067"],"dc:language":["eng"],"dc:subject":["Computer Science"],"dc:title":["Large Scale Information Integration on the Web: Finding, Understanding and Querying Web Databases"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:16Z"}