{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/81604"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/81604","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Integrating Similarity Based Retrieval and Query Refinement in Databases","abstract":"\"With the emergence of many application domains that require imprecise similarity-based access to information, techniques to support such a retrieval paradigm over database systems have emerged as a critical area of research. There are two major areas to this search paradigm. The first one is how to interpret similarity search queries in a relational database context. We address this problem by extending relational operators to natively understand similarity-based retrieval and provide similarity operators that act on user-defined data-types. The second major research area is how to interactively improve the query through user interaction (query refinement). We address this problem by extending some information retrieval and machine learning techniques to affect the interpretation of similarity predicates and the query search condition itself. The result is a new query that better satisfies the users information need. The semantics of this domain favor a \"\"top-k\"\" retrieval approach where we only seek the best matching results for a query. We therefore developed for each of the two areas (similarity search and query refinement) several techniques to efficiently process queries. For similarity queries our query processing algorithms naturally support the top-k retrieval paradigm by returning the answers in order of their similarity to the query. For query refinement, our query processing algorithms reuse the results from previous queries to quickly answer the new refined queries. The algorithms avoid duplicating work done before and perform only the minimum work needed to return the next answer, thus resulting in up to an order of magnitude performance improvement over a naive re-execution of a refined query.\"","abstract_html":"&quot;With the emergence of many application domains that require imprecise similarity-based access to information, techniques to support such a retrieval paradigm over database systems have emerged as a critical area of research. There are two major areas to this search paradigm. The first one is how to interpret similarity search queries in a relational database context. We address this problem by extending relational operators to natively understand similarity-based retrieval and provide similarity operators that act on user-defined data-types. The second major research area is how to interactively improve the query through user interaction (query refinement). We address this problem by extending some information retrieval and machine learning techniques to affect the interpretation of similarity predicates and the query search condition itself. The result is a new query that better satisfies the users information need. The semantics of this domain favor a &quot;&quot;top-k&quot;&quot; retrieval approach where we only seek the best matching results for a query. We therefore developed for each of the two areas (similarity search and query refinement) several techniques to efficiently process queries. For similarity queries our query processing algorithms naturally support the top-k retrieval paradigm by returning the answers in order of their similarity to the query. For query refinement, our query processing algorithms reuse the results from previous queries to quickly answer the new refined queries. The algorithms avoid duplicating work done before and perform only the minimum work needed to return the next answer, thus resulting in up to an order of magnitude performance improvement over a naive re-execution of a refined query.&quot;","abstract_has_math":false,"creators":["Ortega-Binderberger, Michael"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Sharad Mehrotra"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-09-25T20:19:28Z","date_published":"2015-09-25T20:19:28Z","updated_at":"2026-07-22T22:26:16Z","subjects":["Library Science"],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["(MiAaPQ)AAI3070037"],"render_values":[{"text":"(MiAaPQ)AAI3070037","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/2142/81604","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Sharad Mehrotra"]},{"key":"dc:creator","label":"Author","values":["Ortega-Binderberger, Michael"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2015-09-25T20:19:28Z","10000-01-01","2002"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Library Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/81604","(MiAaPQ)AAI3070037"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["\"With the emergence of many application domains that require imprecise similarity-based access to information, techniques to support such a retrieval paradigm over database systems have emerged as a critical area of research. There are two major areas to this search paradigm. The first one is how to interpret similarity search queries in a relational database context. We address this problem by extending relational operators to natively understand similarity-based retrieval and provide similarity operators that act on user-defined data-types. The second major research area is how to interactively improve the query through user interaction (query refinement). We address this problem by extending some information retrieval and machine learning techniques to affect the interpretation of similarity predicates and the query search condition itself. The result is a new query that better satisfies the users information need. The semantics of this domain favor a \"\"top-k\"\" retrieval approach where we only seek the best matching results for a query. We therefore developed for each of the two areas (similarity search and query refinement) several techniques to efficiently process queries. For similarity queries our query processing algorithms naturally support the top-k retrieval paradigm by returning the answers in order of their similarity to the query. For query refinement, our query processing algorithms reuse the results from previous queries to quickly answer the new refined queries. The algorithms avoid duplicating work done before and perform only the minimum work needed to return the next answer, thus resulting in up to an order of magnitude performance improvement over a naive re-execution of a refined query.\"","Made available in DSpace on 2015-09-25T20:19:28Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 3070037.pdf: 10188111 bytes, checksum: 9ee37dedb15cf17ba8d47c96c5f54bbf (MD5) Previous issue date: 2002","Embargo set by: Seth Robbins for item 82885 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","169 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2002."]},{"key":"dc:title","label":"Title","values":["Integrating Similarity Based Retrieval and Query Refinement in Databases"]}]}],"canonical_facts":{"dc:contributor":["Sharad Mehrotra"],"dc:creator":["Ortega-Binderberger, Michael"],"dc:date":["2015-09-25T20:19:28Z","10000-01-01","2002"],"dc:description":["\"With the emergence of many application domains that require imprecise similarity-based access to information, techniques to support such a retrieval paradigm over database systems have emerged as a critical area of research. There are two major areas to this search paradigm. The first one is how to interpret similarity search queries in a relational database context. We address this problem by extending relational operators to natively understand similarity-based retrieval and provide similarity operators that act on user-defined data-types. The second major research area is how to interactively improve the query through user interaction (query refinement). We address this problem by extending some information retrieval and machine learning techniques to affect the interpretation of similarity predicates and the query search condition itself. The result is a new query that better satisfies the users information need. The semantics of this domain favor a \"\"top-k\"\" retrieval approach where we only seek the best matching results for a query. We therefore developed for each of the two areas (similarity search and query refinement) several techniques to efficiently process queries. For similarity queries our query processing algorithms naturally support the top-k retrieval paradigm by returning the answers in order of their similarity to the query. For query refinement, our query processing algorithms reuse the results from previous queries to quickly answer the new refined queries. The algorithms avoid duplicating work done before and perform only the minimum work needed to return the next answer, thus resulting in up to an order of magnitude performance improvement over a naive re-execution of a refined query.\"","Made available in DSpace on 2015-09-25T20:19:28Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 3070037.pdf: 10188111 bytes, checksum: 9ee37dedb15cf17ba8d47c96c5f54bbf (MD5) Previous issue date: 2002","Embargo set by: Seth Robbins for item 82885 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","169 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2002."],"dc:identifier":["http://hdl.handle.net/2142/81604","(MiAaPQ)AAI3070037"],"dc:language":["eng"],"dc:subject":["Library Science"],"dc:title":["Integrating Similarity Based Retrieval and Query Refinement in Databases"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:16Z"}