{"id":{"repo_id":"helsinki","oai_identifier":"oai:helda.helsinki.fi:10138/577810"},"canonical_url":"https://search.dev.ndltd.org/etd/helsinki/oai:helda.helsinki.fi:10138/577810","repository":{"repo_id":"helsinki","name":"University of Helsinki","base_url":"https://helda.helsinki.fi/server/oai/request"},"display":{"title":"Improving Question Answering Systems with Retrieval Augmented Generation","abstract":"Large language models (LLMs) have been proven to be state-of-the-art solutions for many NLP benchmarks. However, LLMs in real applications face many limitations. Although such models are seen to contain real-world knowledge, it is kept implicitly in their parameters that cannot be revised and extended unless expensive additional training is performed. These models can hallucinate by confidently producing human-like texts which might contain misleading information. The knowledge limitation and the tendency to hallucinate cause LLMs to struggle with out-of-domain settings. Furthermore, LLMs lack transparency in that their responses are products of big black-box models. While fine-tuning can mitigate some of these issues, it requires high computing resources. On the other hand, retrieval augmentation has been used to tackle knowledge-intensive tasks and proven by recent studies to be effective when coupled with LLMs. In this thesis, we explore Retrieval-Augmented Generation (RAG), a framework to augment generative LLMs with a neural retriever component, in a domain-specific question answering (QA) task. Empirically, we study how RAG helps LLMs in knowledge-intensive situations and explore design decisions in building a RAG pipeline. Our findings underscore the benefits of RAG in the studied situation by showing that leveraging retrieval augmentation yields significant improvement on QA performance over using a pre-trained LLM alone. Furthermore, incorporating RAG in an LLM-driven QA pipeline results in a QA system that accompanies its predictions with evidence documents, leading to a more trustworthy and grounded AI applications.","abstract_html":"Large language models (LLMs) have been proven to be state-of-the-art solutions for many NLP benchmarks. However, LLMs in real applications face many limitations. Although such models are seen to contain real-world knowledge, it is kept implicitly in their parameters that cannot be revised and extended unless expensive additional training is performed. These models can hallucinate by confidently producing human-like texts which might contain misleading information. The knowledge limitation and the tendency to hallucinate cause LLMs to struggle with out-of-domain settings. Furthermore, LLMs lack transparency in that their responses are products of big black-box models. While fine-tuning can mitigate some of these issues, it requires high computing resources. On the other hand, retrieval augmentation has been used to tackle knowledge-intensive tasks and proven by recent studies to be effective when coupled with LLMs. In this thesis, we explore Retrieval-Augmented Generation (RAG), a framework to augment generative LLMs with a neural retriever component, in a domain-specific question answering (QA) task. Empirically, we study how RAG helps LLMs in knowledge-intensive situations and explore design decisions in building a RAG pipeline. Our findings underscore the benefits of RAG in the studied situation by showing that leveraging retrieval augmentation yields significant improvement on QA performance over using a pre-trained LLM alone. Furthermore, incorporating RAG in an LLM-driven QA pipeline results in a QA system that accompanies its predictions with evidence documents, leading to a more trustworthy and grounded AI applications.","abstract_has_math":false,"creators":["Trangcasanchai, Sathianpong"],"institution":"Helsingin yliopisto","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Helsingin yliopisto, Matemaattis-luonnontieteellinen tiedekunta","University of Helsinki, Faculty of Science","Helsingfors universitet, Matematisk-naturvetenskapliga fakulteten"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024","date_published":"2024","updated_at":"2026-07-27T19:56:04Z","subjects":[],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["URN:NBN:fi:hulib-202406253316"],"render_values":[{"text":"URN:NBN:fi:hulib-202406253316","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/10138/577810","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Helsingin yliopisto, Matemaattis-luonnontieteellinen tiedekunta","University of Helsinki, Faculty of Science","Helsingfors universitet, Matematisk-naturvetenskapliga fakulteten"]},{"key":"dc:creator","label":"Author","values":["Trangcasanchai, Sathianpong"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2024"]},{"key":"dc:publisher","label":"Institution","values":["Helsingin yliopisto","University of Helsinki","Helsingfors universitet"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["URN:NBN:fi:hulib-202406253316","http://hdl.handle.net/10138/577810"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Large language models (LLMs) have been proven to be state-of-the-art solutions for many NLP benchmarks. However, LLMs in real applications face many limitations. Although such models are seen to contain real-world knowledge, it is kept implicitly in their parameters that cannot be revised and extended unless expensive additional training is performed. These models can hallucinate by confidently producing human-like texts which might contain misleading information. The knowledge limitation and the tendency to hallucinate cause LLMs to struggle with out-of-domain settings. Furthermore, LLMs lack transparency in that their responses are products of big black-box models. While fine-tuning can mitigate some of these issues, it requires high computing resources. On the other hand, retrieval augmentation has been used to tackle knowledge-intensive tasks and proven by recent studies to be effective when coupled with LLMs. In this thesis, we explore Retrieval-Augmented Generation (RAG), a framework to augment generative LLMs with a neural retriever component, in a domain-specific question answering (QA) task. Empirically, we study how RAG helps LLMs in knowledge-intensive situations and explore design decisions in building a RAG pipeline. Our findings underscore the benefits of RAG in the studied situation by showing that leveraging retrieval augmentation yields significant improvement on QA performance over using a pre-trained LLM alone. Furthermore, incorporating RAG in an LLM-driven QA pipeline results in a QA system that accompanies its predictions with evidence documents, leading to a more trustworthy and grounded AI applications."]},{"key":"dc:title","label":"Title","values":["Improving Question Answering Systems with Retrieval Augmented Generation"]}]}],"canonical_facts":{"dc:contributor":["Helsingin yliopisto, Matemaattis-luonnontieteellinen tiedekunta","University of Helsinki, Faculty of Science","Helsingfors universitet, Matematisk-naturvetenskapliga fakulteten"],"dc:creator":["Trangcasanchai, Sathianpong"],"dc:date.issued":["2024"],"dc:description.abstract":["Large language models (LLMs) have been proven to be state-of-the-art solutions for many NLP benchmarks. However, LLMs in real applications face many limitations. Although such models are seen to contain real-world knowledge, it is kept implicitly in their parameters that cannot be revised and extended unless expensive additional training is performed. These models can hallucinate by confidently producing human-like texts which might contain misleading information. The knowledge limitation and the tendency to hallucinate cause LLMs to struggle with out-of-domain settings. Furthermore, LLMs lack transparency in that their responses are products of big black-box models. While fine-tuning can mitigate some of these issues, it requires high computing resources. On the other hand, retrieval augmentation has been used to tackle knowledge-intensive tasks and proven by recent studies to be effective when coupled with LLMs. In this thesis, we explore Retrieval-Augmented Generation (RAG), a framework to augment generative LLMs with a neural retriever component, in a domain-specific question answering (QA) task. Empirically, we study how RAG helps LLMs in knowledge-intensive situations and explore design decisions in building a RAG pipeline. Our findings underscore the benefits of RAG in the studied situation by showing that leveraging retrieval augmentation yields significant improvement on QA performance over using a pre-trained LLM alone. Furthermore, incorporating RAG in an LLM-driven QA pipeline results in a QA system that accompanies its predictions with evidence documents, leading to a more trustworthy and grounded AI applications."],"dc:identifier.uri":["URN:NBN:fi:hulib-202406253316","http://hdl.handle.net/10138/577810"],"dc:language.iso":["eng"],"dc:publisher":["Helsingin yliopisto","University of Helsinki","Helsingfors universitet"],"dc:title":["Improving Question Answering Systems with Retrieval Augmented Generation"]},"updated_at":"2026-07-27T19:56:04Z"}