{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/124168"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/124168","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"An systematic evaluation on leading large language models and their factuality investigation as question answering systems","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2024-09-16 without embargo terms","abstract_has_math":false,"creators":["Zheng, Shen"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Chang, Kevin Chen-Chuan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-05","date_published":"2024-05","updated_at":"2026-07-22T22:25:00Z","subjects":["Natural Language Processing","Large Language Model"],"languages":["en","eng"],"rights":["Copyright 2024 Shen Zheng"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/124168","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Chang, Kevin Chen-Chuan"]},{"key":"dc:creator","label":"Author","values":["Zheng, Shen"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-05","2024-04-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Natural Language Processing","Large Language Model"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Shen Zheng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/124168"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","The student, Shen Zheng, accepted the attached license on 2024-04-04 at 16:22.","The student, Shen Zheng, submitted this Thesis for approval on 2024-04-04 at 16:30.","This Thesis was approved for publication on 2024-04-05 at 15:44.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20285 on 2024-09-16 at 00:33:39","This thesis presents a cohesive investigation into the advancements, capabilities, and limitations of current large language models (LLMs), with a focused exploration on their application in question-answering systems. It offers a comprehensive assessment and critique of LLMs, emphasizing their evolution, performance evaluation, and factual accuracy. The first part of the thesis introduces GPT-Fathom, an innovative, open-source evaluation framework designed to systematically assess the performance of over ten leading LLMs, including OpenAI’s legacy models, across a suite of more than twenty benchmarks in seven capability categories, all under uniform testing conditions. This analysis not only tracks the technological progression from GPT-3 to GPT-4, uncovering the incremental benefits of incorporating code data and the effects of various training methodologies like supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF) but also quantifies the so-called alignment tax. The detailed retrospective study provides clarity on the nuanced improvements that contribute to the models’ enhanced reasoning capabilities. Transitioning seamlessly from a broad evaluation of LLM capabilities to a targeted analysis of their application in question-answering systems, the second part scrutinizes the factual reliability of responses provided by models such as ChatGPT. By dissecting failures into categories such as comprehension, factuality, specificity, and inference, this investigation identifies factuality as a predominant area of concern. It further explores the underlying issues of knowledge memorization and recall, suggesting that augmenting LLMs with refined external knowledge bases and optimized recall mechanisms can significantly bolster their factual accuracy. By harmonizing insights from evaluating LLMs’ general capabilities with a deep dive into their performance in question-answering contexts, this thesis elucidates the multifaceted challenges and opportunities facing the development of LLMs. It articulates the necessity for enhanced transparency, accountability, and methodological rigor in the ongoing advancement of LLMs, advocating for strategies that not only improve their intellectual capabilities but also ensure the reliability and truthfulness of their output. Through this integrated analysis, the thesis contributes to a nuanced understanding of LLMs’ current state and charts a forward path for their evolution into more dependable and effective tools in AI-driven applications."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["An systematic evaluation on leading large language models and their factuality investigation as question answering systems"]}]}],"canonical_facts":{"dc:contributor":["Chang, Kevin Chen-Chuan"],"dc:creator":["Zheng, Shen"],"dc:date":["2024-05","2024-04-05"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","The student, Shen Zheng, accepted the attached license on 2024-04-04 at 16:22.","The student, Shen Zheng, submitted this Thesis for approval on 2024-04-04 at 16:30.","This Thesis was approved for publication on 2024-04-05 at 15:44.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20285 on 2024-09-16 at 00:33:39","This thesis presents a cohesive investigation into the advancements, capabilities, and limitations of current large language models (LLMs), with a focused exploration on their application in question-answering systems. It offers a comprehensive assessment and critique of LLMs, emphasizing their evolution, performance evaluation, and factual accuracy. The first part of the thesis introduces GPT-Fathom, an innovative, open-source evaluation framework designed to systematically assess the performance of over ten leading LLMs, including OpenAI’s legacy models, across a suite of more than twenty benchmarks in seven capability categories, all under uniform testing conditions. This analysis not only tracks the technological progression from GPT-3 to GPT-4, uncovering the incremental benefits of incorporating code data and the effects of various training methodologies like supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF) but also quantifies the so-called alignment tax. The detailed retrospective study provides clarity on the nuanced improvements that contribute to the models’ enhanced reasoning capabilities. Transitioning seamlessly from a broad evaluation of LLM capabilities to a targeted analysis of their application in question-answering systems, the second part scrutinizes the factual reliability of responses provided by models such as ChatGPT. By dissecting failures into categories such as comprehension, factuality, specificity, and inference, this investigation identifies factuality as a predominant area of concern. It further explores the underlying issues of knowledge memorization and recall, suggesting that augmenting LLMs with refined external knowledge bases and optimized recall mechanisms can significantly bolster their factual accuracy. By harmonizing insights from evaluating LLMs’ general capabilities with a deep dive into their performance in question-answering contexts, this thesis elucidates the multifaceted challenges and opportunities facing the development of LLMs. It articulates the necessity for enhanced transparency, accountability, and methodological rigor in the ongoing advancement of LLMs, advocating for strategies that not only improve their intellectual capabilities but also ensure the reliability and truthfulness of their output. Through this integrated analysis, the thesis contributes to a nuanced understanding of LLMs’ current state and charts a forward path for their evolution into more dependable and effective tools in AI-driven applications."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/124168"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Shen Zheng"],"dc:subject":["Natural Language Processing","Large Language Model"],"dc:title":["An systematic evaluation on leading large language models and their factuality investigation as question answering systems"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:00Z"}