Back to search

University of Illinois Urbana-Champaign

Diagnostic evaluation of logical reasoning capability of large language models

Abstract

dc:description

Recent advancements in large language models (LLMs) have revolutionized natural language processing, achieving impressive capabilities across diverse tasks; however, critical shortcomings remain, particularly in logical reasoning, characterized by frequent hallucinations, superficial inference patterns, and inconsistencies under linguistic variations. To address these limitations and to pave the foundations for future studies, we introduce a diagnostic evaluation framework utilizing systematically generated synthetic ordering tasks designed to rigorously probe the logical reasoning capacities of LLMs, assessing performance across varying complexities and robustness against controlled perturbations, including equivalent task rephrasings and directional query variations. Evaluating two prominent models, GPT-4.1 and GPTo, at full and reduced ("mini") scales, we found larger models maintain higher accuracy and robustness compared to their mini counterparts but still exhibit declining performance with increasing complexity. Significant sensitivity to linguistic variations was observed, with even strong models displaying inconsistent performance between logically equivalent formulations, and pronounced biases toward certain answer types emerged, indicating reliance on heuristic shortcuts rather than genuine logical deduction. This thesis underscores the necessity for nuanced evaluations beyond simple accuracy metrics, highlighting specific vulnerabilities and robustness limitations, and establishes a foundation for future evaluations aimed at enhancing the reliability and transparency of logical reasoning in LLMs.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois Urbana-Champaign
Year dc:date
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Jiang, Jize
Contributors dc:contributor
  • Zhai, Chengxiang

Subjects

dc:subject × 5

Rights

dc:rights
Statement dc:rights
  • Copyright 2025 Jize Jiang
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/130063

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Jiang, Jize. Diagnostic evaluation of logical reasoning capability of large language models. Thesis thesis, University of Illinois Urbana-Champaign, 2025. https://hdl.handle.net/2142/130063