Back to results

University of Houston

Agentic Framework for Domain-specific RAG Evaluation

Abstract

dc:description.abstract

Large Language Models (LLMs) have revolutionized generation and understanding of textual information and have wide-ranging applications. While the capabilities of LLMs are immediately impressive to any user, extended use quickly reveals problems of inaccurate text generation associated with these models. Retrieval-Augmented Generation (RAG) approaches are used to enhance LLM reliability by grounding responses in knowledge bases, significantly reducing hallucinations. However, RAG performance is highly domain sensitive, necessitating careful tuning of components before deployment for specialized applications. This challenge underscores the critical need for robust evaluation frameworks tailored to domain-specific RAG systems. Existing methods often rely on heuristic-based metrics such as exact match or BLEU scores, which fail to capture deeper semantic reasoning and nuanced understanding. Additionally, these frameworks typically depend on manually curated Question-Answer (QA) datasets, which are often unavailable or insufficient in specialized domains. To address these limitations, we propose an Agentic Framework for Domain-Specific RAG Evaluation. Our approach introduces a synthetic data generation pipeline that simplifies the adaptation of RAG systems to new domains. We incorporate LLM-as-a-judge metrics to enable a more holistic and versatile evaluation of both synthetic datasets and RAG performance. We present MiliQA, a synthetic data set derived from military documents, and compare its quality against public data sets of QA such as Aurelio Mixtral, HuggingFace QA, and WikiEval. To validate our metrics, we benchmark them against human-annotated datasets, including STS-B and SQuAD 2.0. Finally, we demonstrate the applicability of our framework by evaluating key components of RAG systems including embedding models, LLMs, and multiple RAG methodologies, using MiliQA. The results demonstrate the practical value of our proposed framework for guiding the design and optimization of RAG systems for domain-specific applications.

Degree

thesis:*
Name thesis:degree_name
Master of Science
Discipline thesis:degree_discipline
Engineering Data Science
Grantor
University of Houston
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Pham, Quoc Huy 1996-
Advisor dc:contributor.advisor
  • Hoskere, Vedhus
Committee members dc:contributor.committeemember
  • Glennie, Craig L.
  • Bradford, Todd F.
  • Ekhtari, Nima

Subjects

dc:subject × 1

Rights

Language dc:language.iso
English

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/10657/19492
OAI identifier oai:identifier
oai:uh-ir.tdl.org:10657/19492

Chain of custody

source
Harvested from
University of Houston
Base URL
uh-ir.tdl.org/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
related terms
citation

Pham, Quoc Huy 1996-. Agentic Framework for Domain-specific RAG Evaluation. University of Houston, 2025. https://hdl.handle.net/10657/19492