{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132499"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132499","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Robust foundation model for healthcare","abstract":"Clinical data are routinely collected during patient visits, encompassing demographics, diagnoses, laboratory test results, medication prescriptions, medical images, and clinical notes. While there is growing interest in applying deep learning techniques to clinical predictive modeling, developing robust models for healthcare remains challenging due to several key obstacles: (1) data fragmentation, with patient health records scattered across multiple institutions; (2) data missingness, as records frequently contain gaps in both features and labels; (3) distribution shift, where training and testing data often differ in distribution; and (4) complex data schema, given that clinical data comes with varying structures and standardizations. This dissertation presents a comprehensive approach to addressing these challenges through multiple contributions. First, I developed foundational methods for handling healthcare data complexities: MedLink links de-identified patient health records across hospitals by matching health patterns without relying on sensitive patient identifiers; MUSE extends model training to include patients with missing features and labels, leveraging additional training data to improve model performance; SLDG supports the development of clinical predictive models that effectively adapt to domain shifts in target data, ensuring better generalizability; and Llemr introduces a general framework for instruction-tuning large language models (LLMs) to process and interpret clinical data with complex structures. Building upon these methodological contributions, I co-led the development of PyHealth, an open-source Python library that provides a comprehensive framework for deep learning on healthcare data. PyHealth integrates lessons learned from the aforementioned research, offering standardized data processing pipelines, implementations of state-of-the-art healthcare ML models, and evaluation protocols specifically designed for clinical applications. The library provides practical tools for handling complex medical data and supporting various clinical prediction tasks. Together, these contributions, from foundational research methods to practical tools, provide a comprehensive framework for advancing deep learning applications in healthcare, enabling researchers and practitioners to develop more robust and reliable clinical predictive models while addressing the fundamental challenges inherent in medical data.","abstract_html":"Clinical data are routinely collected during patient visits, encompassing demographics, diagnoses, laboratory test results, medication prescriptions, medical images, and clinical notes. While there is growing interest in applying deep learning techniques to clinical predictive modeling, developing robust models for healthcare remains challenging due to several key obstacles: (1) data fragmentation, with patient health records scattered across multiple institutions; (2) data missingness, as records frequently contain gaps in both features and labels; (3) distribution shift, where training and testing data often differ in distribution; and (4) complex data schema, given that clinical data comes with varying structures and standardizations. This dissertation presents a comprehensive approach to addressing these challenges through multiple contributions. First, I developed foundational methods for handling healthcare data complexities: MedLink links de-identified patient health records across hospitals by matching health patterns without relying on sensitive patient identifiers; MUSE extends model training to include patients with missing features and labels, leveraging additional training data to improve model performance; SLDG supports the development of clinical predictive models that effectively adapt to domain shifts in target data, ensuring better generalizability; and Llemr introduces a general framework for instruction-tuning large language models (LLMs) to process and interpret clinical data with complex structures. Building upon these methodological contributions, I co-led the development of PyHealth, an open-source Python library that provides a comprehensive framework for deep learning on healthcare data. PyHealth integrates lessons learned from the aforementioned research, offering standardized data processing pipelines, implementations of state-of-the-art healthcare ML models, and evaluation protocols specifically designed for clinical applications. The library provides practical tools for handling complex medical data and supporting various clinical prediction tasks. Together, these contributions, from foundational research methods to practical tools, provide a comprehensive framework for advancing deep learning applications in healthcare, enabling researchers and practitioners to develop more robust and reliable clinical predictive models while addressing the fundamental challenges inherent in medical data.","abstract_has_math":false,"creators":["Wu, Zhenbang"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Sun, Jimeng","Tong, Hanghang","Zhao, Han","Nalls, Mike","Faghri, Faraz"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["deep learning","healthcare","foundation models"],"languages":["en"],"rights":["Copyright 2025 Zhenbang Wu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132499","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Sun, Jimeng","Tong, Hanghang","Zhao, Han","Nalls, Mike","Faghri, Faraz"]},{"key":"dc:creator","label":"Author","values":["Wu, Zhenbang"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-11-18"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["deep learning","healthcare","foundation models"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Zhenbang Wu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132499"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Clinical data are routinely collected during patient visits, encompassing demographics, diagnoses, laboratory test results, medication prescriptions, medical images, and clinical notes. While there is growing interest in applying deep learning techniques to clinical predictive modeling, developing robust models for healthcare remains challenging due to several key obstacles: (1) data fragmentation, with patient health records scattered across multiple institutions; (2) data missingness, as records frequently contain gaps in both features and labels; (3) distribution shift, where training and testing data often differ in distribution; and (4) complex data schema, given that clinical data comes with varying structures and standardizations. This dissertation presents a comprehensive approach to addressing these challenges through multiple contributions. First, I developed foundational methods for handling healthcare data complexities: MedLink links de-identified patient health records across hospitals by matching health patterns without relying on sensitive patient identifiers; MUSE extends model training to include patients with missing features and labels, leveraging additional training data to improve model performance; SLDG supports the development of clinical predictive models that effectively adapt to domain shifts in target data, ensuring better generalizability; and Llemr introduces a general framework for instruction-tuning large language models (LLMs) to process and interpret clinical data with complex structures. Building upon these methodological contributions, I co-led the development of PyHealth, an open-source Python library that provides a comprehensive framework for deep learning on healthcare data. PyHealth integrates lessons learned from the aforementioned research, offering standardized data processing pipelines, implementations of state-of-the-art healthcare ML models, and evaluation protocols specifically designed for clinical applications. The library provides practical tools for handling complex medical data and supporting various clinical prediction tasks. Together, these contributions, from foundational research methods to practical tools, provide a comprehensive framework for advancing deep learning applications in healthcare, enabling researchers and practitioners to develop more robust and reliable clinical predictive models while addressing the fundamental challenges inherent in medical data.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Zhenbang Wu, accepted the attached license on 2025-11-17 at 22:51.","The student, Zhenbang Wu, submitted this Dissertation for approval on 2025-11-17 at 23:01.","This Dissertation was approved for publication on 2025-11-18 at 13:14.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22880 on 2026-02-19 at 18:24:54"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Robust foundation model for healthcare"]}]}],"canonical_facts":{"dc:contributor":["Sun, Jimeng","Tong, Hanghang","Zhao, Han","Nalls, Mike","Faghri, Faraz"],"dc:creator":["Wu, Zhenbang"],"dc:date":["2025-12","2025-11-18"],"dc:description":["Clinical data are routinely collected during patient visits, encompassing demographics, diagnoses, laboratory test results, medication prescriptions, medical images, and clinical notes. While there is growing interest in applying deep learning techniques to clinical predictive modeling, developing robust models for healthcare remains challenging due to several key obstacles: (1) data fragmentation, with patient health records scattered across multiple institutions; (2) data missingness, as records frequently contain gaps in both features and labels; (3) distribution shift, where training and testing data often differ in distribution; and (4) complex data schema, given that clinical data comes with varying structures and standardizations. This dissertation presents a comprehensive approach to addressing these challenges through multiple contributions. First, I developed foundational methods for handling healthcare data complexities: MedLink links de-identified patient health records across hospitals by matching health patterns without relying on sensitive patient identifiers; MUSE extends model training to include patients with missing features and labels, leveraging additional training data to improve model performance; SLDG supports the development of clinical predictive models that effectively adapt to domain shifts in target data, ensuring better generalizability; and Llemr introduces a general framework for instruction-tuning large language models (LLMs) to process and interpret clinical data with complex structures. Building upon these methodological contributions, I co-led the development of PyHealth, an open-source Python library that provides a comprehensive framework for deep learning on healthcare data. PyHealth integrates lessons learned from the aforementioned research, offering standardized data processing pipelines, implementations of state-of-the-art healthcare ML models, and evaluation protocols specifically designed for clinical applications. The library provides practical tools for handling complex medical data and supporting various clinical prediction tasks. Together, these contributions, from foundational research methods to practical tools, provide a comprehensive framework for advancing deep learning applications in healthcare, enabling researchers and practitioners to develop more robust and reliable clinical predictive models while addressing the fundamental challenges inherent in medical data.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Zhenbang Wu, accepted the attached license on 2025-11-17 at 22:51.","The student, Zhenbang Wu, submitted this Dissertation for approval on 2025-11-17 at 23:01.","This Dissertation was approved for publication on 2025-11-18 at 13:14.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22880 on 2026-02-19 at 18:24:54"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132499"],"dc:language":["en"],"dc:rights":["Copyright 2025 Zhenbang Wu"],"dc:subject":["deep learning","healthcare","foundation models"],"dc:title":["Robust foundation model for healthcare"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}