University of Illinois - Chicago
Structural Robustness of Transformer Models for Clinical Text Summarization on MIMIC-III
Abstract
dc:descriptionTransformer-based models are increasingly used to summarize clinical documents, yet their performance is evaluated exclusively on well-formatted text. In practice, clinical notes undergo structural degradation through hospital mergers, EHR migrations, and copy-paste practices, conditions that no prior study has tested for summarization. This thesis evaluates five models (BART-CNN, BioBART, PEGASUS, Flan-T5-XL, MedAlpaca-7B) across four structural perturbations (control, no headers, no formatting, temporal shuffling) on 1,000 MIMIC-III discharge summaries under de-identified and re-identified conditions, producing 40,000 summaries evaluated against Llama-3 70B silver-standard references. Results reveal widespread fragility: 42/60 perturbation tests and 18/20 cross-condition comparisons reached statistical significance. The study formalizes the Metric-Extraction Paradox; models achieve high ROUGE scores by copying verbatim rather than engaging in clinical synthesis, as both outputs and references share the same source vocabulary. Qualitative analysis of 400 samples identifies six failure modes undetectable by automated metrics, including hallucination, task abandonment, and decoding loops. Key findings challenge prevailing assumptions: Flan-T5-XL, with no medical pre-training, demonstrated the highest robustness, while domain-specific BioBART and MedAlpaca exhibited the most severe failures. Instruction-tuning on diverse tasks outperforms biomedical pre-training for structural resilience.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Mokshit Bharat Surana (24400139)
Subjects
dc:subject × 3Rights
dc:rights- Statement dc:rights
-
- In Copyright
- Open Access after 2028-05-01
Identifiers
dc:identifier.*- DOI dc:identifier
- https://doi.org/10.25417/uic.32995190.v1
- OAI identifier oai:identifier
- oai:figshare.com:article/32995190