University of Tennessee at Chattanooga
A question to query LLM as a pipeline replacement in knowledge graph question answering systems
Abstract
dc:description.abstractKnowledge Graph Question Answering (KGQA) pipelines commonly depend on separate entity and relation predictors before querying the graph, which introduces engineering complexity and costly inference passes over large vocabularies. This thesis presents a drop-in replacement for those modules: a fine-tuned large language model (LLM) that translates a natural-language question directly into an executable SPARQL query. We fine-tune instruction-tuned backbones, Llama-3.1-8B-Instruct and Mistral-7B-Instruct, on paired (question, gold SPARQL) examples, which are formatted through chat templates. As a result, the models can perform single-step query generation. The training and inference pipeline includes a lightweight post-processor that corrects tokenizer-induced spacing artifacts in generated SPARQL, improving exact-match robustness without altering query structure. On a held-out test set, the fine-tuned models achieve 97.9% (Llama) and 94.0% (Mistral) exact-match accuracy for natural-language-to-SPARQL generation, demonstrating that an end-to-end translator can meet or exceed the accuracy of typical multi-module KGQA stacks while substantially simplifying the architecture.
Degree
thesis:*- Grantor dc:publisher
- University of Tennessee at Chattanooga
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Schwartz, Major
- Contributors dc:contributor
-
- Xie, Mengjun
- Sakib, Shahnewaz Karim; Liang, Yu
- College of Engineering and Computer Science
Subjects
dc:subject × 3Rights
dc:rights- Language dc:language
- English, eng
Identifiers
dc:identifier.*- Repository record dc:identifier
- https://scholar.utc.edu/theses/1033
- OAI identifier oai:identifier
- oai:scholar.utc.edu:theses-2221