{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/34551"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/34551","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Systematic optimization of search engines for difficult queries","abstract":"With the advent of Web, text information is being generated across the globe at an unfathomable rate and covering countless topics. This dramatic growth in text information and the increasing number of ways people can utilize it has influenced our daily lives in fundamental and profound ways. The widespread and multi-purposed use of search engines is one example of how transformed our lives have become in relation to text information. Although the vast majority of people's information needs can be served very successfully by current search engines, there is still a considerable number of queries that even the best search engines perform poorly on. In this dissertation, we propose a method of optimizing a search engine to better handle such “difficult queries”, which addresses issues and opportunities at three different stages of an interactive search. Specifically, we propose to improve search quality for difficult queries by: (1) Bridging the vocabulary gap by defining a semantic smoothing language model. A query can be difficult because it does not contain the optimal choice of related terms, or lacks sufficient discriminative terms. In the pre-retrieval stage, using semantic smoothing during query formulation can mitigate the effect of those omissions. (2) Incorporating user negative feedback, i.e., improving a search engine by learning from user feedback in the post-retrieval stage. When a query is so difficult that all the top retrieved results (e.g., top 10) are completely irrelevant, the feedback that a user can provide is solely negative. We propose a generalized optimization framework to learn from feedback on non-relevant documents to prune extensively, but carefully, a large number of non-relevant documents from the top of the ranking list. (3) Balancing priorities when learning from user interaction in order to optimize the whole session. When presenting the search results to the user, there is a tradeoff between promoting those with the highest immediate utility and promoting those with the best potential for collecting feedback information, which can be used to better serve the user over the course of the session (interactive search stage). We frame this tradeoff as a problem of optimizing the diversification of search results, and we propose a machine learning approach that adaptively optimizes each individual user query such that we maximize the overall utility of the entire session. In summary, this dissertation is expected to advance the state of the art of search engines by providing a suite of novel search algorithms, which use the listed approaches to improve a search engine's ability to handle difficult queries.","abstract_html":"With the advent of Web, text information is being generated across the globe at an unfathomable rate and covering countless topics. This dramatic growth in text information and the increasing number of ways people can utilize it has influenced our daily lives in fundamental and profound ways. The widespread and multi-purposed use of search engines is one example of how transformed our lives have become in relation to text information. Although the vast majority of people&#x27;s information needs can be served very successfully by current search engines, there is still a considerable number of queries that even the best search engines perform poorly on. In this dissertation, we propose a method of optimizing a search engine to better handle such “difficult queries”, which addresses issues and opportunities at three different stages of an interactive search. Specifically, we propose to improve search quality for difficult queries by: (1) Bridging the vocabulary gap by defining a semantic smoothing language model. A query can be difficult because it does not contain the optimal choice of related terms, or lacks sufficient discriminative terms. In the pre-retrieval stage, using semantic smoothing during query formulation can mitigate the effect of those omissions. (2) Incorporating user negative feedback, i.e., improving a search engine by learning from user feedback in the post-retrieval stage. When a query is so difficult that all the top retrieved results (e.g., top 10) are completely irrelevant, the feedback that a user can provide is solely negative. We propose a generalized optimization framework to learn from feedback on non-relevant documents to prune extensively, but carefully, a large number of non-relevant documents from the top of the ranking list. (3) Balancing priorities when learning from user interaction in order to optimize the whole session. When presenting the search results to the user, there is a tradeoff between promoting those with the highest immediate utility and promoting those with the best potential for collecting feedback information, which can be used to better serve the user over the course of the session (interactive search stage). We frame this tradeoff as a problem of optimizing the diversification of search results, and we propose a machine learning approach that adaptively optimizes each individual user query such that we maximize the overall utility of the entire session. In summary, this dissertation is expected to advance the state of the art of search engines by providing a suite of novel search algorithms, which use the listed approaches to improve a search engine&#x27;s ability to handle difficult queries.","abstract_has_math":false,"creators":["Karimzadehgan, Maryam"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Zhai, ChengXiang","Han, Jiawei","Efron, Miles J.","Tomkins, Andrew"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2012,"date_issued":"2012-09-18T21:25:15Z","date_published":"2012-09-18T21:25:15Z","updated_at":"2026-07-22T22:25:31Z","subjects":["Information Retrieval","Statistical Translation Language Model","Negative Feedback","Exploration-Exploitation Optimization"],"languages":["en"],"rights":["Copyright 2012 Maryam Karimzadehgan"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/34551","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhai, ChengXiang","Han, Jiawei","Efron, Miles J.","Tomkins, Andrew"]},{"key":"dc:creator","label":"Author","values":["Karimzadehgan, Maryam"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2012-09-18T21:25:15Z","2014-09-18T10:00:47Z","2012-08"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Information Retrieval","Statistical Translation Language Model","Negative Feedback","Exploration-Exploitation Optimization"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2012 Maryam Karimzadehgan"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/34551"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["With the advent of Web, text information is being generated across the globe at an unfathomable rate and covering countless topics. This dramatic growth in text information and the increasing number of ways people can utilize it has influenced our daily lives in fundamental and profound ways. The widespread and multi-purposed use of search engines is one example of how transformed our lives have become in relation to text information. Although the vast majority of people's information needs can be served very successfully by current search engines, there is still a considerable number of queries that even the best search engines perform poorly on. In this dissertation, we propose a method of optimizing a search engine to better handle such “difficult queries”, which addresses issues and opportunities at three different stages of an interactive search. Specifically, we propose to improve search quality for difficult queries by: (1) Bridging the vocabulary gap by defining a semantic smoothing language model. A query can be difficult because it does not contain the optimal choice of related terms, or lacks sufficient discriminative terms. In the pre-retrieval stage, using semantic smoothing during query formulation can mitigate the effect of those omissions. (2) Incorporating user negative feedback, i.e., improving a search engine by learning from user feedback in the post-retrieval stage. When a query is so difficult that all the top retrieved results (e.g., top 10) are completely irrelevant, the feedback that a user can provide is solely negative. We propose a generalized optimization framework to learn from feedback on non-relevant documents to prune extensively, but carefully, a large number of non-relevant documents from the top of the ranking list. (3) Balancing priorities when learning from user interaction in order to optimize the whole session. When presenting the search results to the user, there is a tradeoff between promoting those with the highest immediate utility and promoting those with the best potential for collecting feedback information, which can be used to better serve the user over the course of the session (interactive search stage). We frame this tradeoff as a problem of optimizing the diversification of search results, and we propose a machine learning approach that adaptively optimizes each individual user query such that we maximize the overall utility of the entire session. In summary, this dissertation is expected to advance the state of the art of search engines by providing a suite of novel search algorithms, which use the listed approaches to improve a search engine's ability to handle difficult queries.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2012-06-21T14:42:46Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 karimzadehgan_maryam.pdf: 890304 bytes, checksum: 1f205dccdee7bcef2378f5b5c20be11a (MD5)","Made available in DSpace on 2012-09-18T21:25:15Z (GMT). No. of bitstreams: 2 karimzadehgan_maryam.pdf: 890304 bytes, checksum: 1f205dccdee7bcef2378f5b5c20be11a (MD5) license.txt: 4071 bytes, checksum: cd63edc736e0b9e167689c1aa96d2952 (MD5)","Restriction data tranferred 2014-07-01T11:35:48-05:00 Original Data Group with Access Administrator Release Date: 2014-09-18 16:27:16 UTC Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Item marked as restricted to the 'Administrator' Group (id=1) by Seth Robbins (srobbins@illinois.edu) on 2012-09-18T21:27:27Z Item is restricted until 2014-09-18T21:27:16Z","Limited Restriction Lifted for Item 34825 on 2014-09-18T10:00:47Z."]},{"key":"dc:title","label":"Title","values":["Systematic optimization of search engines for difficult queries"]}]}],"canonical_facts":{"dc:contributor":["Zhai, ChengXiang","Han, Jiawei","Efron, Miles J.","Tomkins, Andrew"],"dc:creator":["Karimzadehgan, Maryam"],"dc:date":["2012-09-18T21:25:15Z","2014-09-18T10:00:47Z","2012-08"],"dc:description":["With the advent of Web, text information is being generated across the globe at an unfathomable rate and covering countless topics. This dramatic growth in text information and the increasing number of ways people can utilize it has influenced our daily lives in fundamental and profound ways. The widespread and multi-purposed use of search engines is one example of how transformed our lives have become in relation to text information. Although the vast majority of people's information needs can be served very successfully by current search engines, there is still a considerable number of queries that even the best search engines perform poorly on. In this dissertation, we propose a method of optimizing a search engine to better handle such “difficult queries”, which addresses issues and opportunities at three different stages of an interactive search. Specifically, we propose to improve search quality for difficult queries by: (1) Bridging the vocabulary gap by defining a semantic smoothing language model. A query can be difficult because it does not contain the optimal choice of related terms, or lacks sufficient discriminative terms. In the pre-retrieval stage, using semantic smoothing during query formulation can mitigate the effect of those omissions. (2) Incorporating user negative feedback, i.e., improving a search engine by learning from user feedback in the post-retrieval stage. When a query is so difficult that all the top retrieved results (e.g., top 10) are completely irrelevant, the feedback that a user can provide is solely negative. We propose a generalized optimization framework to learn from feedback on non-relevant documents to prune extensively, but carefully, a large number of non-relevant documents from the top of the ranking list. (3) Balancing priorities when learning from user interaction in order to optimize the whole session. When presenting the search results to the user, there is a tradeoff between promoting those with the highest immediate utility and promoting those with the best potential for collecting feedback information, which can be used to better serve the user over the course of the session (interactive search stage). We frame this tradeoff as a problem of optimizing the diversification of search results, and we propose a machine learning approach that adaptively optimizes each individual user query such that we maximize the overall utility of the entire session. In summary, this dissertation is expected to advance the state of the art of search engines by providing a suite of novel search algorithms, which use the listed approaches to improve a search engine's ability to handle difficult queries.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2012-06-21T14:42:46Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 karimzadehgan_maryam.pdf: 890304 bytes, checksum: 1f205dccdee7bcef2378f5b5c20be11a (MD5)","Made available in DSpace on 2012-09-18T21:25:15Z (GMT). No. of bitstreams: 2 karimzadehgan_maryam.pdf: 890304 bytes, checksum: 1f205dccdee7bcef2378f5b5c20be11a (MD5) license.txt: 4071 bytes, checksum: cd63edc736e0b9e167689c1aa96d2952 (MD5)","Restriction data tranferred 2014-07-01T11:35:48-05:00 Original Data Group with Access Administrator Release Date: 2014-09-18 16:27:16 UTC Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Item marked as restricted to the 'Administrator' Group (id=1) by Seth Robbins (srobbins@illinois.edu) on 2012-09-18T21:27:27Z Item is restricted until 2014-09-18T21:27:16Z","Limited Restriction Lifted for Item 34825 on 2014-09-18T10:00:47Z."],"dc:identifier":["http://hdl.handle.net/2142/34551"],"dc:language":["en"],"dc:rights":["Copyright 2012 Maryam Karimzadehgan"],"dc:subject":["Information Retrieval","Statistical Translation Language Model","Negative Feedback","Exploration-Exploitation Optimization"],"dc:title":["Systematic optimization of search engines for difficult queries"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:31Z"}