{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/130063"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/130063","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Diagnostic evaluation of logical reasoning capability of large language models","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-08-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2027-08-01","abstract_has_math":false,"creators":["Jiang, Jize"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Zhai, Chengxiang"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-07-22","date_published":"2025-07-22","updated_at":"2026-07-22T22:25:06Z","subjects":["Natural Language Processing","Machine Learning","Large Language Model","Logic Reasoning","Model Evaluation"],"languages":["en","eng"],"rights":["Copyright 2025 Jize Jiang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/130063","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhai, Chengxiang"]},{"key":"dc:creator","label":"Author","values":["Jiang, Jize"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-07-22","2025-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Natural Language Processing","Machine Learning","Large Language Model","Logic Reasoning","Model Evaluation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Jize Jiang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/130063"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-08-01","The student, Jize Jiang, accepted the attached license on 2025-07-22 at 15:56.","The student, Jize Jiang, submitted this Thesis for approval on 2025-07-22 at 16:09.","This Thesis was approved for publication on 2025-07-22 at 16:46.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22698 on 2025-10-21 at 10:06:19","Recent advancements in large language models (LLMs) have revolutionized natural language processing, achieving impressive capabilities across diverse tasks; however, critical shortcomings remain, particularly in logical reasoning, characterized by frequent hallucinations, superficial inference patterns, and inconsistencies under linguistic variations. To address these limitations and to pave the foundations for future studies, we introduce a diagnostic evaluation framework utilizing systematically generated synthetic ordering tasks designed to rigorously probe the logical reasoning capacities of LLMs, assessing performance across varying complexities and robustness against controlled perturbations, including equivalent task rephrasings and directional query variations. Evaluating two prominent models, GPT-4.1 and GPTo, at full and reduced (\"mini\") scales, we found larger models maintain higher accuracy and robustness compared to their mini counterparts but still exhibit declining performance with increasing complexity. Significant sensitivity to linguistic variations was observed, with even strong models displaying inconsistent performance between logically equivalent formulations, and pronounced biases toward certain answer types emerged, indicating reliance on heuristic shortcuts rather than genuine logical deduction. This thesis underscores the necessity for nuanced evaluations beyond simple accuracy metrics, highlighting specific vulnerabilities and robustness limitations, and establishes a foundation for future evaluations aimed at enhancing the reliability and transparency of logical reasoning in LLMs."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Diagnostic evaluation of logical reasoning capability of large language models"]}]}],"canonical_facts":{"dc:contributor":["Zhai, Chengxiang"],"dc:creator":["Jiang, Jize"],"dc:date":["2025-07-22","2025-08"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-08-01","The student, Jize Jiang, accepted the attached license on 2025-07-22 at 15:56.","The student, Jize Jiang, submitted this Thesis for approval on 2025-07-22 at 16:09.","This Thesis was approved for publication on 2025-07-22 at 16:46.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22698 on 2025-10-21 at 10:06:19","Recent advancements in large language models (LLMs) have revolutionized natural language processing, achieving impressive capabilities across diverse tasks; however, critical shortcomings remain, particularly in logical reasoning, characterized by frequent hallucinations, superficial inference patterns, and inconsistencies under linguistic variations. To address these limitations and to pave the foundations for future studies, we introduce a diagnostic evaluation framework utilizing systematically generated synthetic ordering tasks designed to rigorously probe the logical reasoning capacities of LLMs, assessing performance across varying complexities and robustness against controlled perturbations, including equivalent task rephrasings and directional query variations. Evaluating two prominent models, GPT-4.1 and GPTo, at full and reduced (\"mini\") scales, we found larger models maintain higher accuracy and robustness compared to their mini counterparts but still exhibit declining performance with increasing complexity. Significant sensitivity to linguistic variations was observed, with even strong models displaying inconsistent performance between logically equivalent formulations, and pronounced biases toward certain answer types emerged, indicating reliance on heuristic shortcuts rather than genuine logical deduction. This thesis underscores the necessity for nuanced evaluations beyond simple accuracy metrics, highlighting specific vulnerabilities and robustness limitations, and establishes a foundation for future evaluations aimed at enhancing the reliability and transparency of logical reasoning in LLMs."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/130063"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Jize Jiang"],"dc:subject":["Natural Language Processing","Machine Learning","Large Language Model","Logic Reasoning","Model Evaluation"],"dc:title":["Diagnostic evaluation of logical reasoning capability of large language models"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:06Z"}