{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132602"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132602","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"An end-to-end benchmarking framework for retrieval-augmented generation systems","abstract":"As Large Language Models (LLMs) transition from experimental prototypes to production-grade services, Retrieval-Augmented Generation (RAG) has emerged as the de facto paradigm for mitigating hallucinations and incorporating up-to-date knowledge. While the accuracy of RAG systems has been extensively studied, the system performance—specifically regarding throughput, latency, and memory efficiency at scale—remains largely unexplored. However, current RAG benchmarking efforts are predominantly accuracy-centric, focusing on metrics like precision and recall while neglecting the implications of the underlying retrieval infrastructure and generation bottlenecks. Consequently, developers face significant challenges in navigating the complex trade-offs between vector database configurations, retrieval strategies, and generative model parameters. This thesis presents a RAG-based AI system benchmarking framework (RASB) for characterizing the system performance of RAG pipelines. To enable a holistic evaluation, RASB decouples the RAG workflow into modular components—embedding, indexing, retrieval, and generation—allowing for fine-grained analysis of each stage. We rethink the evaluation methodology by shifting the focus from pure answer quality to system efficiency, exploring multiple dimensions such as varying batch sizes, vector database index types, and embedding dimensions. RASB provides a testbed that supports modular RAG pipelines with major vector databases and LLM backends, automating the collection of performance metrics that include end-to-end throughput, GPU memory consumption, and context recall. To evaluate diverse usage scenarios, RASB integrates a configurable workload generator that drives experiments using both real-world and synthetic datasets. We demonstrate RASB’s capability through a comprehensive set of experiments conducted on popular Vector Databases and LLM backends.","abstract_html":"As Large Language Models (LLMs) transition from experimental prototypes to production-grade services, Retrieval-Augmented Generation (RAG) has emerged as the de facto paradigm for mitigating hallucinations and incorporating up-to-date knowledge. While the accuracy of RAG systems has been extensively studied, the system performance—specifically regarding throughput, latency, and memory efficiency at scale—remains largely unexplored. However, current RAG benchmarking efforts are predominantly accuracy-centric, focusing on metrics like precision and recall while neglecting the implications of the underlying retrieval infrastructure and generation bottlenecks. Consequently, developers face significant challenges in navigating the complex trade-offs between vector database configurations, retrieval strategies, and generative model parameters. This thesis presents a RAG-based AI system benchmarking framework (RASB) for characterizing the system performance of RAG pipelines. To enable a holistic evaluation, RASB decouples the RAG workflow into modular components—embedding, indexing, retrieval, and generation—allowing for fine-grained analysis of each stage. We rethink the evaluation methodology by shifting the focus from pure answer quality to system efficiency, exploring multiple dimensions such as varying batch sizes, vector database index types, and embedding dimensions. RASB provides a testbed that supports modular RAG pipelines with major vector databases and LLM backends, automating the collection of performance metrics that include end-to-end throughput, GPU memory consumption, and context recall. To evaluate diverse usage scenarios, RASB integrates a configurable workload generator that drives experiments using both real-world and synthetic datasets. We demonstrate RASB’s capability through a comprehensive set of experiments conducted on popular Vector Databases and LLM backends.","abstract_has_math":false,"creators":["Xu, Yuan"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Huang, Jian"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["Retrieval-Augmented Generation (RAG)","Benchmarking framework","AI system performance"],"languages":["en"],"rights":["Copyright 2025 Yuan Xu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132602","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Huang, Jian"]},{"key":"dc:creator","label":"Author","values":["Xu, Yuan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-12-11"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Retrieval-Augmented Generation (RAG)","Benchmarking framework","AI system performance"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Yuan Xu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132602"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["As Large Language Models (LLMs) transition from experimental prototypes to production-grade services, Retrieval-Augmented Generation (RAG) has emerged as the de facto paradigm for mitigating hallucinations and incorporating up-to-date knowledge. While the accuracy of RAG systems has been extensively studied, the system performance—specifically regarding throughput, latency, and memory efficiency at scale—remains largely unexplored. However, current RAG benchmarking efforts are predominantly accuracy-centric, focusing on metrics like precision and recall while neglecting the implications of the underlying retrieval infrastructure and generation bottlenecks. Consequently, developers face significant challenges in navigating the complex trade-offs between vector database configurations, retrieval strategies, and generative model parameters. This thesis presents a RAG-based AI system benchmarking framework (RASB) for characterizing the system performance of RAG pipelines. To enable a holistic evaluation, RASB decouples the RAG workflow into modular components—embedding, indexing, retrieval, and generation—allowing for fine-grained analysis of each stage. We rethink the evaluation methodology by shifting the focus from pure answer quality to system efficiency, exploring multiple dimensions such as varying batch sizes, vector database index types, and embedding dimensions. RASB provides a testbed that supports modular RAG pipelines with major vector databases and LLM backends, automating the collection of performance metrics that include end-to-end throughput, GPU memory consumption, and context recall. To evaluate diverse usage scenarios, RASB integrates a configurable workload generator that drives experiments using both real-world and synthetic datasets. We demonstrate RASB’s capability through a comprehensive set of experiments conducted on popular Vector Databases and LLM backends.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Yuan Xu, accepted the attached license on 2025-12-11 at 16:16.","The student, Yuan Xu, submitted this Thesis for approval on 2025-12-11 at 17:07.","This Thesis was approved for publication on 2025-12-11 at 17:17.","DSpace SAF Submission Ingestion Package generated from Vireo submission #23138 on 2026-02-19 at 18:30:17"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["An end-to-end benchmarking framework for retrieval-augmented generation systems"]}]}],"canonical_facts":{"dc:contributor":["Huang, Jian"],"dc:creator":["Xu, Yuan"],"dc:date":["2025-12","2025-12-11"],"dc:description":["As Large Language Models (LLMs) transition from experimental prototypes to production-grade services, Retrieval-Augmented Generation (RAG) has emerged as the de facto paradigm for mitigating hallucinations and incorporating up-to-date knowledge. While the accuracy of RAG systems has been extensively studied, the system performance—specifically regarding throughput, latency, and memory efficiency at scale—remains largely unexplored. However, current RAG benchmarking efforts are predominantly accuracy-centric, focusing on metrics like precision and recall while neglecting the implications of the underlying retrieval infrastructure and generation bottlenecks. Consequently, developers face significant challenges in navigating the complex trade-offs between vector database configurations, retrieval strategies, and generative model parameters. This thesis presents a RAG-based AI system benchmarking framework (RASB) for characterizing the system performance of RAG pipelines. To enable a holistic evaluation, RASB decouples the RAG workflow into modular components—embedding, indexing, retrieval, and generation—allowing for fine-grained analysis of each stage. We rethink the evaluation methodology by shifting the focus from pure answer quality to system efficiency, exploring multiple dimensions such as varying batch sizes, vector database index types, and embedding dimensions. RASB provides a testbed that supports modular RAG pipelines with major vector databases and LLM backends, automating the collection of performance metrics that include end-to-end throughput, GPU memory consumption, and context recall. To evaluate diverse usage scenarios, RASB integrates a configurable workload generator that drives experiments using both real-world and synthetic datasets. We demonstrate RASB’s capability through a comprehensive set of experiments conducted on popular Vector Databases and LLM backends.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Yuan Xu, accepted the attached license on 2025-12-11 at 16:16.","The student, Yuan Xu, submitted this Thesis for approval on 2025-12-11 at 17:07.","This Thesis was approved for publication on 2025-12-11 at 17:17.","DSpace SAF Submission Ingestion Package generated from Vireo submission #23138 on 2026-02-19 at 18:30:17"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132602"],"dc:language":["en"],"dc:rights":["Copyright 2025 Yuan Xu"],"dc:subject":["Retrieval-Augmented Generation (RAG)","Benchmarking framework","AI system performance"],"dc:title":["An end-to-end benchmarking framework for retrieval-augmented generation systems"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}