{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/81787"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/81787","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Efficient Data Integration: Automation, Collaboration, and Relaxation","abstract":"While the previous two directions reduce integration costs by improving the performance of automatic tools (either by improvements to the tool itself, or by leveraging users to boost tool accuracy), the last direction explored in this thesis attacks data integration costs at their foundation---rigidity. The current data integration system model imposes a very rigid structure on its components and the data that is passed between components. For example, wrappers are responsible for extracting precise structured data, allowing traditional structured query processing techniques to compute the query result. However, my third direction explores our ability to relax these assumptions, thereby allowing us to answer queries without suffering unnecessary costs required in the traditional model (e.g., building full-fledged wrappers). In this thesis I investigate this idea within the context of supporting one-time, on-the-fly queries over distributed Web data. I develop and evaluate SLIC, a system that allows a user to quickly pose SQL queries over multiple sources (after only some minimal preprocessing), obtain initial results, then iterate with the system to get increasingly better results. The fundamental idea is to learn only as much structure as necessary to answer a given query. Extensive experiments on real-world domains show that for many practical queries SLIC is significantly faster than current methods, thus providing a promising first step toward a principled solution for lazy, on-the-fly integration of Web data, and hopefully sparking interest in our potential to remove some of the fundamental costs inherent in the traditional integration system model.","abstract_html":"While the previous two directions reduce integration costs by improving the performance of automatic tools (either by improvements to the tool itself, or by leveraging users to boost tool accuracy), the last direction explored in this thesis attacks data integration costs at their foundation---rigidity. The current data integration system model imposes a very rigid structure on its components and the data that is passed between components. For example, wrappers are responsible for extracting precise structured data, allowing traditional structured query processing techniques to compute the query result. However, my third direction explores our ability to relax these assumptions, thereby allowing us to answer queries without suffering unnecessary costs required in the traditional model (e.g., building full-fledged wrappers). In this thesis I investigate this idea within the context of supporting one-time, on-the-fly queries over distributed Web data. I develop and evaluate SLIC, a system that allows a user to quickly pose SQL queries over multiple sources (after only some minimal preprocessing), obtain initial results, then iterate with the system to get increasingly better results. The fundamental idea is to learn only as much structure as necessary to answer a given query. Extensive experiments on real-world domains show that for many practical queries SLIC is significantly faster than current methods, thus providing a promising first step toward a principled solution for lazy, on-the-fly integration of Web data, and hopefully sparking interest in our potential to remove some of the fundamental costs inherent in the traditional integration system model.","abstract_has_math":false,"creators":["McCann, Robert Lee"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Doan, AnHai"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-09-25T20:20:27Z","date_published":"2015-09-25T20:20:27Z","updated_at":"2026-07-22T22:26:16Z","subjects":["Computer Science"],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["(MiAaPQ)AAI3290314"],"render_values":[{"text":"(MiAaPQ)AAI3290314","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/2142/81787","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Doan, AnHai"]},{"key":"dc:creator","label":"Author","values":["McCann, Robert Lee"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2015-09-25T20:20:27Z","10000-01-01","2007"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/81787","(MiAaPQ)AAI3290314"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["While the previous two directions reduce integration costs by improving the performance of automatic tools (either by improvements to the tool itself, or by leveraging users to boost tool accuracy), the last direction explored in this thesis attacks data integration costs at their foundation---rigidity. The current data integration system model imposes a very rigid structure on its components and the data that is passed between components. For example, wrappers are responsible for extracting precise structured data, allowing traditional structured query processing techniques to compute the query result. However, my third direction explores our ability to relax these assumptions, thereby allowing us to answer queries without suffering unnecessary costs required in the traditional model (e.g., building full-fledged wrappers). In this thesis I investigate this idea within the context of supporting one-time, on-the-fly queries over distributed Web data. I develop and evaluate SLIC, a system that allows a user to quickly pose SQL queries over multiple sources (after only some minimal preprocessing), obtain initial results, then iterate with the system to get increasingly better results. The fundamental idea is to learn only as much structure as necessary to answer a given query. Extensive experiments on real-world domains show that for many practical queries SLIC is significantly faster than current methods, thus providing a promising first step toward a principled solution for lazy, on-the-fly integration of Web data, and hopefully sparking interest in our potential to remove some of the fundamental costs inherent in the traditional integration system model.","Made available in DSpace on 2015-09-25T20:20:27Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 3290314.pdf: 4340764 bytes, checksum: 41ed28ca7e8c7b8887a06ec578522709 (MD5) Previous issue date: 2007","Embargo set by: Seth Robbins for item 83068 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","121 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2007."]},{"key":"dc:title","label":"Title","values":["Efficient Data Integration: Automation, Collaboration, and Relaxation"]}]}],"canonical_facts":{"dc:contributor":["Doan, AnHai"],"dc:creator":["McCann, Robert Lee"],"dc:date":["2015-09-25T20:20:27Z","10000-01-01","2007"],"dc:description":["While the previous two directions reduce integration costs by improving the performance of automatic tools (either by improvements to the tool itself, or by leveraging users to boost tool accuracy), the last direction explored in this thesis attacks data integration costs at their foundation---rigidity. The current data integration system model imposes a very rigid structure on its components and the data that is passed between components. For example, wrappers are responsible for extracting precise structured data, allowing traditional structured query processing techniques to compute the query result. However, my third direction explores our ability to relax these assumptions, thereby allowing us to answer queries without suffering unnecessary costs required in the traditional model (e.g., building full-fledged wrappers). In this thesis I investigate this idea within the context of supporting one-time, on-the-fly queries over distributed Web data. I develop and evaluate SLIC, a system that allows a user to quickly pose SQL queries over multiple sources (after only some minimal preprocessing), obtain initial results, then iterate with the system to get increasingly better results. The fundamental idea is to learn only as much structure as necessary to answer a given query. Extensive experiments on real-world domains show that for many practical queries SLIC is significantly faster than current methods, thus providing a promising first step toward a principled solution for lazy, on-the-fly integration of Web data, and hopefully sparking interest in our potential to remove some of the fundamental costs inherent in the traditional integration system model.","Made available in DSpace on 2015-09-25T20:20:27Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 3290314.pdf: 4340764 bytes, checksum: 41ed28ca7e8c7b8887a06ec578522709 (MD5) Previous issue date: 2007","Embargo set by: Seth Robbins for item 83068 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","121 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2007."],"dc:identifier":["http://hdl.handle.net/2142/81787","(MiAaPQ)AAI3290314"],"dc:language":["eng"],"dc:subject":["Computer Science"],"dc:title":["Efficient Data Integration: Automation, Collaboration, and Relaxation"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:16Z"}