Back to search

University of Illinois - Chicago

Towards Interpretable, Equitable, and Safe Language Model Applications in Healthcare and Beyond

Abstract

dc:description

Over the past few years, large language models (LLMs) have taken the landscape of natural language processing (NLP) by storm, delivering promising results across various tasks. However, their deployment in high-stakes settings remains challenging due to concerns about interpretability, fairness, and safety. This thesis investigates these three interconnected dimensions of trustworthy language model applications, with a particular focus on healthcare. We begin by focusing on interpretable language-model-based conversational agents in resource-constrained healthcare settings and propose modularized, neuro-symbolic dialogue architectures for health coaching, maintaining interpretability and control while requiring minimal annotation. Having established interpretable architectures, we examine whether LLMs can be trusted to perform equitably across patient demographics. We find that state-of-the-art LLMs consistently predict less favorable outcomes for African American patients, make stereotypical linguistic assumptions about race, and fail to apply medical knowledge equitably despite strong performance on textbook benchmarks. More concerningly, we find that LLMs exhibit biases in domains where demographics are not priors for predicting outcomes. Introducing the concept of Veracity Bias, we show that LLMs systematically associate solution correctness with demographic groups in mathematics, coding, reasoning, and essay assessment tasks. The persistence of such biases despite alignment suggests that current LLM alignment can be superficial, leading to reasoning failures that extend beyond fairness. We demonstrate this through the Fallacy Failure Attack, which exploits LLMs' inability to generate convincingly false information and cause models to leak truthful, harmful content while bypassing safeguard mechanisms. These findings have significant implications for deploying language models in high-stakes applications. As LLMs become increasingly integrated into decision-making systems across healthcare, education, and beyond, ensuring their trustworthiness is not merely a technical challenge but a societal imperative. Moving forward, the research community should prioritize not just improving model capabilities but fundamentally rethinking how we build AI systems that are interpretable, fair, and secure.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Yue Zhou (71752)

Subjects

dc:subject × 2

Rights

dc:rights
Statement dc:rights
  • In Copyright

Identifiers

dc:identifier.*
OAI identifier oai:identifier
oai:figshare.com:article/32991953

Chain of custody

source
Harvested from
University of Illinois - Chicago
Base URL
api.figshare.com/v2/oai
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Yue Zhou (71752). Towards Interpretable, Equitable, and Safe Language Model Applications in Healthcare and Beyond. 2026. https://doi.org/10.25417/uic.32991953.v1