University of Illinois Urbana-Champaign
Malice, inequality, instability, or ignorance? Disentangling the mechanisms of LLM unfairness
Abstract
dc:descriptionEnsuring fairness in large language models (LLMs) is critical as these models are increasingly deployed in sensitive domains. Traditional fairness metrics typically report a single scalar score, which conflates distinct sources of model failure and obscures underlying biases. In this work, we propose a Hierarchical Bias-Variance Decomposition framework—termed BDSU—that decomposes total discrimination risk into four interpretable components: Bias (systematic global error), Disparity (group-level variance), Sensitivity (context-level variance), and Uncertainty (stochastic or token-level variance). By applying the law of total variance recursively, BDSU provides a principled method to quantify and separate these failure modes, aligning each with ethical and reliability priorities. We further introduce a conditional micro-diagnosis to evaluate fairness at the group level, enabling fine-grained auditing and targeted interventions. Our theoretical framework lays the foundation for more transparent, actionable, and robust evaluation of LLM fairness, highlighting the distinct mechanisms by which models may perpetuate bias or exhibit instability.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois Urbana-Champaign
- Year dc:date
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Yang, Ke
- Contributors dc:contributor
-
- Zhai, ChengXiang
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Copyright © 2025 Ke Yang. All rights reserved.
- Language dc:language
- en
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/132559
- OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/132559