{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129865"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129865","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"An agentic approach to information seeking for knowledge acquisition","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-20 without embargo terms","abstract_has_math":false,"creators":["Gangi Reddy, Revanth"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Ji, Heng","Zhai, ChengXiang","Hakkani-Tur, Dilek","Das, Dipanjan","Wen-tau Yih, Scott"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-07-14","date_published":"2025-07-14","updated_at":"2026-07-22T22:25:06Z","subjects":["Information Seeking","Agentic Search","Retrieval And Ranking","Knowledge Extraction"],"languages":["en","eng"],"rights":["Copyright 2025 Revanth Gangi Reddy"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129865","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ji, Heng","Zhai, ChengXiang","Hakkani-Tur, Dilek","Das, Dipanjan","Wen-tau Yih, Scott"]},{"key":"dc:creator","label":"Author","values":["Gangi Reddy, Revanth"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-07-14","2025-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Information Seeking","Agentic Search","Retrieval And Ranking","Knowledge Extraction"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Revanth Gangi Reddy"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129865"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","The student, Revanth Gangi Reddy, accepted the attached license on 2025-07-11 at 11:51.","The student, Revanth Gangi Reddy, submitted this Dissertation for approval on 2025-07-11 at 11:56.","This Dissertation was approved for publication on 2025-07-14 at 09:48.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22510 on 2025-10-20 at 16:57:47","The digital age has created an information paradox: while access to data is unprecedented, our ability to synthesize it into actionable knowledge is increasingly overwhelmed. Modern search systems, optimized for simple, linear tasks, are unsuited for the multi-source aggregation and iterative reasoning that defines human-like information seeking. Further, current web agents fail at the breadth-first exploration required to answer complex queries. This thesis bridges these critical gaps by introducing a new paradigm for automated information seeking, moving beyond simple retrieval to a more holistic and human-like approach to knowledge acquisition. The foundational contribution is a modular, agent-based framework that decomposes the complex information-seeking process into specialized Navigator, Extractor, and Aggregator components. Guided by a feedback loop, our proposed approach intelligently explores the web and aggregates information across a variety of information access settings, enabling a more human-like, exploratory approach to information gathering. The true utility of this general framework is demonstrated through its adaptation to high-impact, specialized domains. To combat information latency in crowd-sourced knowledge bases, we introduce an agentic system that automates Wikipedia updates by monitoring online sources and generating edit suggestions for human review. For the software development domain, where code’s unique semantics challenge traditional search, we leverage large-scale, curated code data to train specialized ranking models that achieve state-of-the-art performance in code retrieval and software issue localization. An effective information-seeking system is fundamentally dependent on its core ranking and retrieval engine. This thesis addresses three critical bottlenecks in modern search systems. To overcome the limitation of fixed-granularity ranking, we introduce an approach that leverages multi-vector embeddings for any-granularity ranking, allowing flexible retrieval from passages down to atomic propositions without separate indexing. To resolve the efficiency-effectiveness trade-off of powerful but slow listwise LLM rerankers, we propose ranking with a single token output, boosting inference latency by up to 50% while improving ranking accuracy. Finally, to break the recall ceiling imposed by initial retrieval, we introduce an inference-time feedback mechanism that uses the reranker's output to refine the retriever's query representations for improving search recall. Beyond finding documents, true understanding requires extracting precise, structured knowledge. This dissertation introduces novel methods for fine-grained knowledge extraction involving claim detection and cross-media question answering. We introduce a zero-shot framework that uses question answering to extract not only factual claims from news but also their critical attributes, such as the claim source and subject topic. For contexts where information is split across modalities, we drive progress in complex question answering that requires grounding entities between images and text for cross-media reasoning. These advancements culminate in a human-centered system that combines the principles of information-seeking and knowledge acquisition to address the complex needs of intelligence analysis. By automating the generation of structured, verifiable, and event-driven situation reports, we provide a powerful demonstration of how the research in this thesis can empower humans to focus on high-level strategic thinking. Together, these contributions provide a robust foundation for the next generation of information-seeking systems, that can enable humans to more effectively navigate and make sense of our complex, information-rich world."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["An agentic approach to information seeking for knowledge acquisition"]}]}],"canonical_facts":{"dc:contributor":["Ji, Heng","Zhai, ChengXiang","Hakkani-Tur, Dilek","Das, Dipanjan","Wen-tau Yih, Scott"],"dc:creator":["Gangi Reddy, Revanth"],"dc:date":["2025-07-14","2025-08"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","The student, Revanth Gangi Reddy, accepted the attached license on 2025-07-11 at 11:51.","The student, Revanth Gangi Reddy, submitted this Dissertation for approval on 2025-07-11 at 11:56.","This Dissertation was approved for publication on 2025-07-14 at 09:48.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22510 on 2025-10-20 at 16:57:47","The digital age has created an information paradox: while access to data is unprecedented, our ability to synthesize it into actionable knowledge is increasingly overwhelmed. Modern search systems, optimized for simple, linear tasks, are unsuited for the multi-source aggregation and iterative reasoning that defines human-like information seeking. Further, current web agents fail at the breadth-first exploration required to answer complex queries. This thesis bridges these critical gaps by introducing a new paradigm for automated information seeking, moving beyond simple retrieval to a more holistic and human-like approach to knowledge acquisition. The foundational contribution is a modular, agent-based framework that decomposes the complex information-seeking process into specialized Navigator, Extractor, and Aggregator components. Guided by a feedback loop, our proposed approach intelligently explores the web and aggregates information across a variety of information access settings, enabling a more human-like, exploratory approach to information gathering. The true utility of this general framework is demonstrated through its adaptation to high-impact, specialized domains. To combat information latency in crowd-sourced knowledge bases, we introduce an agentic system that automates Wikipedia updates by monitoring online sources and generating edit suggestions for human review. For the software development domain, where code’s unique semantics challenge traditional search, we leverage large-scale, curated code data to train specialized ranking models that achieve state-of-the-art performance in code retrieval and software issue localization. An effective information-seeking system is fundamentally dependent on its core ranking and retrieval engine. This thesis addresses three critical bottlenecks in modern search systems. To overcome the limitation of fixed-granularity ranking, we introduce an approach that leverages multi-vector embeddings for any-granularity ranking, allowing flexible retrieval from passages down to atomic propositions without separate indexing. To resolve the efficiency-effectiveness trade-off of powerful but slow listwise LLM rerankers, we propose ranking with a single token output, boosting inference latency by up to 50% while improving ranking accuracy. Finally, to break the recall ceiling imposed by initial retrieval, we introduce an inference-time feedback mechanism that uses the reranker's output to refine the retriever's query representations for improving search recall. Beyond finding documents, true understanding requires extracting precise, structured knowledge. This dissertation introduces novel methods for fine-grained knowledge extraction involving claim detection and cross-media question answering. We introduce a zero-shot framework that uses question answering to extract not only factual claims from news but also their critical attributes, such as the claim source and subject topic. For contexts where information is split across modalities, we drive progress in complex question answering that requires grounding entities between images and text for cross-media reasoning. These advancements culminate in a human-centered system that combines the principles of information-seeking and knowledge acquisition to address the complex needs of intelligence analysis. By automating the generation of structured, verifiable, and event-driven situation reports, we provide a powerful demonstration of how the research in this thesis can empower humans to focus on high-level strategic thinking. Together, these contributions provide a robust foundation for the next generation of information-seeking systems, that can enable humans to more effectively navigate and make sense of our complex, information-rich world."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129865"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Revanth Gangi Reddy"],"dc:subject":["Information Seeking","Agentic Search","Retrieval And Ranking","Knowledge Extraction"],"dc:title":["An agentic approach to information seeking for knowledge acquisition"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:06Z"}