{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132493"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132493","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Anticipating and resolving ad hoc information needs during discourse","abstract":"The ability to access information is the foundation of a functional society, but this access is sometimes hindered by the challenges of resolving ad hoc information needs - those that arise spontaneously while a user is exposed to new information (e.g., while a student listening to a lecture). Traditionally, these needs are manually resolved, which is often limited by personal or environmental constraints. This thesis addresses this problem by studying methods for anticipating and automatically resolving likely ad hoc information needs, contributing five main studies motivated by gaps in the existing literature. First, we study how to generate likely student questions using online, asynchronous lecture video transcripts in low-data settings. After curating a dataset of lecture transcripts and asked questions, we use low-data techniques to train generative language models, and find that pre-training with search engine queries leads to more precise questions but continuous prefix tuning offers mixed results. Second, we investigate how retrieval models perform for anticipating and resolving ad hoc information needs using different pre-search contexts. We evaluate the models using a constructed dataset of webpages mentioned in online discussion threads, and find that the models struggle to proactively anticipate needed information. Third, we build on the aforementioned contributions by developing TextData, a system demonstration guided by our first two contributions that provides users with predicted questions or search results based on highlighted context. Fourth, we build InstInfo, another system demonstration that suggests research papers to academic conference presentation attendees in real-time. We conduct a small user study with InstInfo, and find that the needs of conference attendees are mostly centered around content from the paper being presented, and that timing is critical for user utility. Fifth, based on this user study, we develop the RTPR (Real-Time Proactive Retrieval) framework, which is a suite of evaluation metrics for evaluating real-time proactive retrieval systems. We then apply these metrics to retrieval models trained using a newly-created dataset (ProPres, a refined academic conference setting), and find significant limitations for existing retrieval models to provide user utility in settings of real-time proactive retrieval. Collectively, the contribution of this thesis is the systematic exploration of how to anticipate and resolve ad hoc information needs across various settings, with the goal of laying the groundwork for future proactive retrieval systems.","abstract_html":"The ability to access information is the foundation of a functional society, but this access is sometimes hindered by the challenges of resolving ad hoc information needs - those that arise spontaneously while a user is exposed to new information (e.g., while a student listening to a lecture). Traditionally, these needs are manually resolved, which is often limited by personal or environmental constraints. This thesis addresses this problem by studying methods for anticipating and automatically resolving likely ad hoc information needs, contributing five main studies motivated by gaps in the existing literature. First, we study how to generate likely student questions using online, asynchronous lecture video transcripts in low-data settings. After curating a dataset of lecture transcripts and asked questions, we use low-data techniques to train generative language models, and find that pre-training with search engine queries leads to more precise questions but continuous prefix tuning offers mixed results. Second, we investigate how retrieval models perform for anticipating and resolving ad hoc information needs using different pre-search contexts. We evaluate the models using a constructed dataset of webpages mentioned in online discussion threads, and find that the models struggle to proactively anticipate needed information. Third, we build on the aforementioned contributions by developing TextData, a system demonstration guided by our first two contributions that provides users with predicted questions or search results based on highlighted context. Fourth, we build InstInfo, another system demonstration that suggests research papers to academic conference presentation attendees in real-time. We conduct a small user study with InstInfo, and find that the needs of conference attendees are mostly centered around content from the paper being presented, and that timing is critical for user utility. Fifth, based on this user study, we develop the RTPR (Real-Time Proactive Retrieval) framework, which is a suite of evaluation metrics for evaluating real-time proactive retrieval systems. We then apply these metrics to retrieval models trained using a newly-created dataset (ProPres, a refined academic conference setting), and find significant limitations for existing retrieval models to provide user utility in settings of real-time proactive retrieval. Collectively, the contribution of this thesis is the systematic exploration of how to anticipate and resolve ad hoc information needs across various settings, with the goal of laying the groundwork for future proactive retrieval systems.","abstract_has_math":false,"creators":["Ros, Kevin"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Zhai, ChengXiang","Chen-Chuan Chang, Kevin","Huang, Yun","Deng, Yu"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["information retrieval","proactive","search engines","information needs"],"languages":["en"],"rights":["Copyright 2025 Kevin Ros"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132493","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhai, ChengXiang","Chen-Chuan Chang, Kevin","Huang, Yun","Deng, Yu"]},{"key":"dc:creator","label":"Author","values":["Ros, Kevin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-11-14"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["information retrieval","proactive","search engines","information needs"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Kevin Ros"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132493"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The ability to access information is the foundation of a functional society, but this access is sometimes hindered by the challenges of resolving ad hoc information needs - those that arise spontaneously while a user is exposed to new information (e.g., while a student listening to a lecture). Traditionally, these needs are manually resolved, which is often limited by personal or environmental constraints. This thesis addresses this problem by studying methods for anticipating and automatically resolving likely ad hoc information needs, contributing five main studies motivated by gaps in the existing literature. First, we study how to generate likely student questions using online, asynchronous lecture video transcripts in low-data settings. After curating a dataset of lecture transcripts and asked questions, we use low-data techniques to train generative language models, and find that pre-training with search engine queries leads to more precise questions but continuous prefix tuning offers mixed results. Second, we investigate how retrieval models perform for anticipating and resolving ad hoc information needs using different pre-search contexts. We evaluate the models using a constructed dataset of webpages mentioned in online discussion threads, and find that the models struggle to proactively anticipate needed information. Third, we build on the aforementioned contributions by developing TextData, a system demonstration guided by our first two contributions that provides users with predicted questions or search results based on highlighted context. Fourth, we build InstInfo, another system demonstration that suggests research papers to academic conference presentation attendees in real-time. We conduct a small user study with InstInfo, and find that the needs of conference attendees are mostly centered around content from the paper being presented, and that timing is critical for user utility. Fifth, based on this user study, we develop the RTPR (Real-Time Proactive Retrieval) framework, which is a suite of evaluation metrics for evaluating real-time proactive retrieval systems. We then apply these metrics to retrieval models trained using a newly-created dataset (ProPres, a refined academic conference setting), and find significant limitations for existing retrieval models to provide user utility in settings of real-time proactive retrieval. Collectively, the contribution of this thesis is the systematic exploration of how to anticipate and resolve ad hoc information needs across various settings, with the goal of laying the groundwork for future proactive retrieval systems.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Kevin Ros, accepted the attached license on 2025-11-14 at 10:06.","The student, Kevin Ros, submitted this Dissertation for approval on 2025-11-14 at 10:19.","This Dissertation was approved for publication on 2025-11-14 at 15:42.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22869 on 2026-02-19 at 18:24:41"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Anticipating and resolving ad hoc information needs during discourse"]}]}],"canonical_facts":{"dc:contributor":["Zhai, ChengXiang","Chen-Chuan Chang, Kevin","Huang, Yun","Deng, Yu"],"dc:creator":["Ros, Kevin"],"dc:date":["2025-12","2025-11-14"],"dc:description":["The ability to access information is the foundation of a functional society, but this access is sometimes hindered by the challenges of resolving ad hoc information needs - those that arise spontaneously while a user is exposed to new information (e.g., while a student listening to a lecture). Traditionally, these needs are manually resolved, which is often limited by personal or environmental constraints. This thesis addresses this problem by studying methods for anticipating and automatically resolving likely ad hoc information needs, contributing five main studies motivated by gaps in the existing literature. First, we study how to generate likely student questions using online, asynchronous lecture video transcripts in low-data settings. After curating a dataset of lecture transcripts and asked questions, we use low-data techniques to train generative language models, and find that pre-training with search engine queries leads to more precise questions but continuous prefix tuning offers mixed results. Second, we investigate how retrieval models perform for anticipating and resolving ad hoc information needs using different pre-search contexts. We evaluate the models using a constructed dataset of webpages mentioned in online discussion threads, and find that the models struggle to proactively anticipate needed information. Third, we build on the aforementioned contributions by developing TextData, a system demonstration guided by our first two contributions that provides users with predicted questions or search results based on highlighted context. Fourth, we build InstInfo, another system demonstration that suggests research papers to academic conference presentation attendees in real-time. We conduct a small user study with InstInfo, and find that the needs of conference attendees are mostly centered around content from the paper being presented, and that timing is critical for user utility. Fifth, based on this user study, we develop the RTPR (Real-Time Proactive Retrieval) framework, which is a suite of evaluation metrics for evaluating real-time proactive retrieval systems. We then apply these metrics to retrieval models trained using a newly-created dataset (ProPres, a refined academic conference setting), and find significant limitations for existing retrieval models to provide user utility in settings of real-time proactive retrieval. Collectively, the contribution of this thesis is the systematic exploration of how to anticipate and resolve ad hoc information needs across various settings, with the goal of laying the groundwork for future proactive retrieval systems.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Kevin Ros, accepted the attached license on 2025-11-14 at 10:06.","The student, Kevin Ros, submitted this Dissertation for approval on 2025-11-14 at 10:19.","This Dissertation was approved for publication on 2025-11-14 at 15:42.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22869 on 2026-02-19 at 18:24:41"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132493"],"dc:language":["en"],"dc:rights":["Copyright 2025 Kevin Ros"],"dc:subject":["information retrieval","proactive","search engines","information needs"],"dc:title":["Anticipating and resolving ad hoc information needs during discourse"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}