{"id":{"repo_id":"uic","oai_identifier":"oai:figshare.com:article/32995190"},"canonical_url":"https://search.dev.ndltd.org/etd/uic/oai:figshare.com:article/32995190","repository":{"repo_id":"uic","name":"University of Illinois - Chicago","base_url":"https://api.figshare.com/v2/oai"},"display":{"title":"Structural Robustness of Transformer Models for Clinical Text Summarization on MIMIC-III","abstract":"Transformer-based models are increasingly used to summarize clinical documents, yet their performance is evaluated exclusively on well-formatted text. In practice, clinical notes undergo structural degradation through hospital mergers, EHR migrations, and copy-paste practices, conditions that no prior study has tested for summarization. This thesis evaluates five models (BART-CNN, BioBART, PEGASUS, Flan-T5-XL, MedAlpaca-7B) across four structural perturbations (control, no headers, no formatting, temporal shuffling) on 1,000 MIMIC-III discharge summaries under de-identified and re-identified conditions, producing 40,000 summaries evaluated against Llama-3 70B silver-standard references. Results reveal widespread fragility: 42/60 perturbation tests and 18/20 cross-condition comparisons reached statistical significance. The study formalizes the Metric-Extraction Paradox; models achieve high ROUGE scores by copying verbatim rather than engaging in clinical synthesis, as both outputs and references share the same source vocabulary. Qualitative analysis of 400 samples identifies six failure modes undetectable by automated metrics, including hallucination, task abandonment, and decoding loops. Key findings challenge prevailing assumptions: Flan-T5-XL, with no medical pre-training, demonstrated the highest robustness, while domain-specific BioBART and MedAlpaca exhibited the most severe failures. Instruction-tuning on diverse tasks outperforms biomedical pre-training for structural resilience.","abstract_html":"Transformer-based models are increasingly used to summarize clinical documents, yet their performance is evaluated exclusively on well-formatted text. In practice, clinical notes undergo structural degradation through hospital mergers, EHR migrations, and copy-paste practices, conditions that no prior study has tested for summarization. This thesis evaluates five models (BART-CNN, BioBART, PEGASUS, Flan-T5-XL, MedAlpaca-7B) across four structural perturbations (control, no headers, no formatting, temporal shuffling) on 1,000 MIMIC-III discharge summaries under de-identified and re-identified conditions, producing 40,000 summaries evaluated against Llama-3 70B silver-standard references. Results reveal widespread fragility: 42/60 perturbation tests and 18/20 cross-condition comparisons reached statistical significance. The study formalizes the Metric-Extraction Paradox; models achieve high ROUGE scores by copying verbatim rather than engaging in clinical synthesis, as both outputs and references share the same source vocabulary. Qualitative analysis of 400 samples identifies six failure modes undetectable by automated metrics, including hallucination, task abandonment, and decoding loops. Key findings challenge prevailing assumptions: Flan-T5-XL, with no medical pre-training, demonstrated the highest robustness, while domain-specific BioBART and MedAlpaca exhibited the most severe failures. Instruction-tuning on diverse tasks outperforms biomedical pre-training for structural resilience.","abstract_has_math":false,"creators":["Mokshit Bharat Surana (24400139)"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2026,"date_issued":"2026-05-01T00:00:00Z","date_published":"2026-05-01T00:00:00Z","updated_at":"2026-07-27T21:33:51Z","subjects":["Computer Science","Natural Language Processing","Health Informatics"],"languages":[],"rights":["In Copyright","Open Access after 2028-05-01"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.25417/uic.32995190.v1","outbound_label":"DOI","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Mokshit Bharat Surana (24400139)"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2026-05-01T00:00:00Z"]},{"key":"dc:relation","label":"Dc Relation","values":["https://figshare.com/articles/thesis/Structural_Robustness_of_Transformer_Models_for_Clinical_Text_Summarization_on_MIMIC-III/32995190"]},{"key":"dc:type","label":"Dc Type","values":["Text","Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Science","Natural Language Processing","Health Informatics"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["In Copyright","Open Access after 2028-05-01"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["10.25417/uic.32995190.v1"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Transformer-based models are increasingly used to summarize clinical documents, yet their performance is evaluated exclusively on well-formatted text. In practice, clinical notes undergo structural degradation through hospital mergers, EHR migrations, and copy-paste practices, conditions that no prior study has tested for summarization. This thesis evaluates five models (BART-CNN, BioBART, PEGASUS, Flan-T5-XL, MedAlpaca-7B) across four structural perturbations (control, no headers, no formatting, temporal shuffling) on 1,000 MIMIC-III discharge summaries under de-identified and re-identified conditions, producing 40,000 summaries evaluated against Llama-3 70B silver-standard references. Results reveal widespread fragility: 42/60 perturbation tests and 18/20 cross-condition comparisons reached statistical significance. The study formalizes the Metric-Extraction Paradox; models achieve high ROUGE scores by copying verbatim rather than engaging in clinical synthesis, as both outputs and references share the same source vocabulary. Qualitative analysis of 400 samples identifies six failure modes undetectable by automated metrics, including hallucination, task abandonment, and decoding loops. Key findings challenge prevailing assumptions: Flan-T5-XL, with no medical pre-training, demonstrated the highest robustness, while domain-specific BioBART and MedAlpaca exhibited the most severe failures. Instruction-tuning on diverse tasks outperforms biomedical pre-training for structural resilience."]},{"key":"dc:title","label":"Title","values":["Structural Robustness of Transformer Models for Clinical Text Summarization on MIMIC-III"]}]}],"canonical_facts":{"dc:creator":["Mokshit Bharat Surana (24400139)"],"dc:date":["2026-05-01T00:00:00Z"],"dc:description":["Transformer-based models are increasingly used to summarize clinical documents, yet their performance is evaluated exclusively on well-formatted text. In practice, clinical notes undergo structural degradation through hospital mergers, EHR migrations, and copy-paste practices, conditions that no prior study has tested for summarization. This thesis evaluates five models (BART-CNN, BioBART, PEGASUS, Flan-T5-XL, MedAlpaca-7B) across four structural perturbations (control, no headers, no formatting, temporal shuffling) on 1,000 MIMIC-III discharge summaries under de-identified and re-identified conditions, producing 40,000 summaries evaluated against Llama-3 70B silver-standard references. Results reveal widespread fragility: 42/60 perturbation tests and 18/20 cross-condition comparisons reached statistical significance. The study formalizes the Metric-Extraction Paradox; models achieve high ROUGE scores by copying verbatim rather than engaging in clinical synthesis, as both outputs and references share the same source vocabulary. Qualitative analysis of 400 samples identifies six failure modes undetectable by automated metrics, including hallucination, task abandonment, and decoding loops. Key findings challenge prevailing assumptions: Flan-T5-XL, with no medical pre-training, demonstrated the highest robustness, while domain-specific BioBART and MedAlpaca exhibited the most severe failures. Instruction-tuning on diverse tasks outperforms biomedical pre-training for structural resilience."],"dc:identifier":["10.25417/uic.32995190.v1"],"dc:relation":["https://figshare.com/articles/thesis/Structural_Robustness_of_Transformer_Models_for_Clinical_Text_Summarization_on_MIMIC-III/32995190"],"dc:rights":["In Copyright","Open Access after 2028-05-01"],"dc:subject":["Computer Science","Natural Language Processing","Health Informatics"],"dc:title":["Structural Robustness of Transformer Models for Clinical Text Summarization on MIMIC-III"],"dc:type":["Text","Thesis"]},"updated_at":"2026-07-27T21:33:51Z"}