University of Illinois - Chicago
Towards Interpretable, Equitable, and Safe Language Model Applications in Healthcare and Beyond
Abstract
dc:descriptionOver the past few years, large language models (LLMs) have taken the landscape of natural language processing (NLP) by storm, delivering promising results across various tasks. However, their deployment in high-stakes settings remains challenging due to concerns about interpretability, fairness, and safety. This thesis investigates these three interconnected dimensions of trustworthy language model applications, with a particular focus on healthcare. We begin by focusing on interpretable language-model-based conversational agents in resource-constrained healthcare settings and propose modularized, neuro-symbolic dialogue architectures for health coaching, maintaining interpretability and control while requiring minimal annotation. Having established interpretable architectures, we examine whether LLMs can be trusted to perform equitably across patient demographics. We find that state-of-the-art LLMs consistently predict less favorable outcomes for African American patients, make stereotypical linguistic assumptions about race, and fail to apply medical knowledge equitably despite strong performance on textbook benchmarks. More concerningly, we find that LLMs exhibit biases in domains where demographics are not priors for predicting outcomes. Introducing the concept of Veracity Bias, we show that LLMs systematically associate solution correctness with demographic groups in mathematics, coding, reasoning, and essay assessment tasks. The persistence of such biases despite alignment suggests that current LLM alignment can be superficial, leading to reasoning failures that extend beyond fairness. We demonstrate this through the Fallacy Failure Attack, which exploits LLMs' inability to generate convincingly false information and cause models to leak truthful, harmful content while bypassing safeguard mechanisms. These findings have significant implications for deploying language models in high-stakes applications. As LLMs become increasingly integrated into decision-making systems across healthcare, education, and beyond, ensuring their trustworthiness is not merely a technical challenge but a societal imperative. Moving forward, the research community should prioritize not just improving model capabilities but fundamentally rethinking how we build AI systems that are interpretable, fair, and secure.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Yue Zhou (71752)
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- In Copyright
Identifiers
dc:identifier.*- DOI dc:identifier
- https://doi.org/10.25417/uic.32991953.v1
- OAI identifier oai:identifier
- oai:figshare.com:article/32991953