{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129272"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129272","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Improving speculative retrieval-augmented generation via verifier scoring","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_has_math":false,"creators":["Chen, Shixin"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Kindratenko, Volodymyr"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-04-30","date_published":"2025-04-30","updated_at":"2026-07-22T22:25:04Z","subjects":["Speculative Decoding","RAG","LLMs"],"languages":["en","eng"],"rights":["Copyright 2025 Shixin Chen"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129272","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Kindratenko, Volodymyr"]},{"key":"dc:creator","label":"Author","values":["Chen, Shixin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-04-30","2025-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Speculative Decoding","RAG","LLMs"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Shixin Chen"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129272"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Shixin Chen, accepted the attached license on 2025-04-29 at 01:58.","The student, Shixin Chen, submitted this Thesis for approval on 2025-04-29 at 02:07.","This Thesis was approved for publication on 2025-04-30 at 11:03.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22062 on 2025-10-19 at 18:11:13","The rapid progress of large language models (LLMs) has revolutionized various natural language processing (NLP) tasks, including text generation, question answering, and summarization. However, the computational cost and latency associated with LLMs remain significant challenges, particularly in real-time applications. What's more, the heavy fine-tuning cost for each downstream task is also unbearable in the preparation phase. To address these issues, researchers have explored various techniques to accelerate both inference and fine-tuning without compromising the quality of the generated text. One of the recent breakthroughs is Speculative Decoding. It leverages a smaller, faster \"draft\" model to propose multiple candidate drafts that are then verified in parallel by a larger, more generative, and more accurate \"verification\" model. This method significantly reduces the number of calls to the larger model, thereby speeding up the inference process. Also, the drafter model is much easier to fine-tune than larger models. In this thesis, we are inspired by an innovative extension of the speculative decoding framework by integrating Retrieval-Augmented Generation (RAG). RAG enhances the capabilities of language models by retrieving relevant documents from a knowledge base to inform the generation process. Our research represents a comprehensive analysis Speculative RAG, aims to further optimize the inference pipeline by categorizing the retrieved documents and feeding them to a smaller drafter model, which generates multiple draft continuations in parallel. These drafts are then evaluated by a larger, general-purpose verifier model to select the most accurate and contextually appropriate continuation."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Improving speculative retrieval-augmented generation via verifier scoring"]}]}],"canonical_facts":{"dc:contributor":["Kindratenko, Volodymyr"],"dc:creator":["Chen, Shixin"],"dc:date":["2025-04-30","2025-05"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Shixin Chen, accepted the attached license on 2025-04-29 at 01:58.","The student, Shixin Chen, submitted this Thesis for approval on 2025-04-29 at 02:07.","This Thesis was approved for publication on 2025-04-30 at 11:03.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22062 on 2025-10-19 at 18:11:13","The rapid progress of large language models (LLMs) has revolutionized various natural language processing (NLP) tasks, including text generation, question answering, and summarization. However, the computational cost and latency associated with LLMs remain significant challenges, particularly in real-time applications. What's more, the heavy fine-tuning cost for each downstream task is also unbearable in the preparation phase. To address these issues, researchers have explored various techniques to accelerate both inference and fine-tuning without compromising the quality of the generated text. One of the recent breakthroughs is Speculative Decoding. It leverages a smaller, faster \"draft\" model to propose multiple candidate drafts that are then verified in parallel by a larger, more generative, and more accurate \"verification\" model. This method significantly reduces the number of calls to the larger model, thereby speeding up the inference process. Also, the drafter model is much easier to fine-tune than larger models. In this thesis, we are inspired by an innovative extension of the speculative decoding framework by integrating Retrieval-Augmented Generation (RAG). RAG enhances the capabilities of language models by retrieving relevant documents from a knowledge base to inform the generation process. Our research represents a comprehensive analysis Speculative RAG, aims to further optimize the inference pipeline by categorizing the retrieved documents and feeding them to a smaller drafter model, which generates multiple draft continuations in parallel. These drafts are then evaluated by a larger, general-purpose verifier model to select the most accurate and contextually appropriate continuation."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129272"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Shixin Chen"],"dc:subject":["Speculative Decoding","RAG","LLMs"],"dc:title":["Improving speculative retrieval-augmented generation via verifier scoring"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:04Z"}