{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129278"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129278","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"The Illinois Retrieval Benchmark : A scalable framework for characterizing retrieval-augmented generation via automated fact-checking","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_has_math":false,"creators":["Taleka, Bhagyashree"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hwu, Wen-mei"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-05-07","date_published":"2025-05-07","updated_at":"2026-07-22T22:25:04Z","subjects":["Retrieval-augmented Generation","Large Language Models","Automated Fact Checking","Evaluation Benchmark","Scalable Dataset"],"languages":["en","eng"],"rights":["Copyright 2025 Bhagyashree Taleka"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129278","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hwu, Wen-mei"]},{"key":"dc:creator","label":"Author","values":["Taleka, Bhagyashree"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-05-07","2025-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Retrieval-augmented Generation","Large Language Models","Automated Fact Checking","Evaluation Benchmark","Scalable Dataset"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Bhagyashree Taleka"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129278"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Bhagyashree Taleka, accepted the attached license on 2025-04-30 at 09:00.","The student, Bhagyashree Taleka, submitted this Thesis for approval on 2025-04-30 at 09:05.","This Thesis was approved for publication on 2025-05-07 at 14:42.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22071 on 2025-10-19 at 18:11:14","Retrieval-Augmented Generation (RAG) combines a retrieval module with a generation module to enhance response accuracy by incorporating external knowledge. To evaluate the robustness, faithfulness, and generation capabilities of RAG systems, numerous evaluation benchmarks have been proposed. However, these benchmarks either focus on evaluating specific modules using large-scale data that provides ground truth documents and query embeddings without access to raw text, or they evaluate RAG outputs using large language models (LLMs) on small datasets, without assessing the behavior or interactions of individual components within the pipeline. Crucially, there exists no benchmark that connects all components of the RAG pipeline and enables systematic end-to-end analysis across retrieval, reranking, and generation. To address these gaps, we propose the Illinois Retrieval Benchmark (IRB), an end-to-end RAG characterization framework designed to enable scalable and customizable dataset generation across diverse domains and retrieval contexts. Our methodology can be used to generate scalable datasets for developing RAG pipelines and evaluating them using pertinent metrics such as context relevance, answer relevance, and answer faithfulness. We also offer a novel approach to generating ground-truth data for user queries. As part of this thesis, we developed a small dataset consisting of 57,671 unique facts, 117,234 unique questions based on these facts, and 2,462,190 document chunks in the vector database to demonstrate the usability and efficacy of our benchmark in evaluating RAG systems across diverse domains and retrieval contexts. Our experiments show that while rerankers improve answer quality with reduced context size, the effectiveness of reranking depends on the quality of base retrieval. Our ablation studies show that existing retrieval methods exhibit significant deficiencies. We also highlight the impact of document-level noise on RAG performance and emphasize the need for accurate and semantically aligned context, particularly when dealing with emerging knowledge domains. Ultimately, we argue that improving LLM performance requires enhancing the retrieval process or continuously fine-tuning models with updated data."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["The Illinois Retrieval Benchmark : A scalable framework for characterizing retrieval-augmented generation via automated fact-checking"]}]}],"canonical_facts":{"dc:contributor":["Hwu, Wen-mei"],"dc:creator":["Taleka, Bhagyashree"],"dc:date":["2025-05-07","2025-05"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Bhagyashree Taleka, accepted the attached license on 2025-04-30 at 09:00.","The student, Bhagyashree Taleka, submitted this Thesis for approval on 2025-04-30 at 09:05.","This Thesis was approved for publication on 2025-05-07 at 14:42.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22071 on 2025-10-19 at 18:11:14","Retrieval-Augmented Generation (RAG) combines a retrieval module with a generation module to enhance response accuracy by incorporating external knowledge. To evaluate the robustness, faithfulness, and generation capabilities of RAG systems, numerous evaluation benchmarks have been proposed. However, these benchmarks either focus on evaluating specific modules using large-scale data that provides ground truth documents and query embeddings without access to raw text, or they evaluate RAG outputs using large language models (LLMs) on small datasets, without assessing the behavior or interactions of individual components within the pipeline. Crucially, there exists no benchmark that connects all components of the RAG pipeline and enables systematic end-to-end analysis across retrieval, reranking, and generation. To address these gaps, we propose the Illinois Retrieval Benchmark (IRB), an end-to-end RAG characterization framework designed to enable scalable and customizable dataset generation across diverse domains and retrieval contexts. Our methodology can be used to generate scalable datasets for developing RAG pipelines and evaluating them using pertinent metrics such as context relevance, answer relevance, and answer faithfulness. We also offer a novel approach to generating ground-truth data for user queries. As part of this thesis, we developed a small dataset consisting of 57,671 unique facts, 117,234 unique questions based on these facts, and 2,462,190 document chunks in the vector database to demonstrate the usability and efficacy of our benchmark in evaluating RAG systems across diverse domains and retrieval contexts. Our experiments show that while rerankers improve answer quality with reduced context size, the effectiveness of reranking depends on the quality of base retrieval. Our ablation studies show that existing retrieval methods exhibit significant deficiencies. We also highlight the impact of document-level noise on RAG performance and emphasize the need for accurate and semantically aligned context, particularly when dealing with emerging knowledge domains. Ultimately, we argue that improving LLM performance requires enhancing the retrieval process or continuously fine-tuning models with updated data."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129278"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Bhagyashree Taleka"],"dc:subject":["Retrieval-augmented Generation","Large Language Models","Automated Fact Checking","Evaluation Benchmark","Scalable Dataset"],"dc:title":["The Illinois Retrieval Benchmark : A scalable framework for characterizing retrieval-augmented generation via automated fact-checking"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:04Z"}