{"id":{"repo_id":"uic","oai_identifier":"oai:figshare.com:article/32994962"},"canonical_url":"https://search.dev.ndltd.org/etd/uic/oai:figshare.com:article/32994962","repository":{"repo_id":"uic","name":"University of Illinois - Chicago","base_url":"https://api.figshare.com/v2/oai"},"display":{"title":"Advancing Precision in Prenatal Depression Prediction: Leveraging EMRs and Enhancing ML-Model Fairness","abstract":"Machine learning models for predicting perinatal depression (PND) from electronic medical records (EMRs) show promise for early intervention, yet significant performance disparities (model bias) across sociodemographic groups limit their clinical utility for diverse populations. Current EMR-based approaches remain unexplored for low-income minority populations and fail to incorporate neighborhood-level information (e.g., pollution, access to healthy food, proximity to hospitals) that is well-known to affect human health. Further, existing computation methods to reduce model bias are designed to correct one attribute at a time (e.g., race, age, socioeconomic status), overlooking how combinations of patient attributes create unique patterns of prediction disparities (e.g., sex and race). This dissertation develops methodological approaches to enhance prediction fairness while maintaining clinical utility in predominantly low-income women of color at an urban academic medical center. Machine learning models were developed to predict early pregnancy depression while systematically assessing model bias using EMRs from individuals receiving obstetric care at UIHealth, where we mostly serve low-income women of color. Our models had worse performance for Black and Hispanic women despite their majority representation the data. Integration of neighborhood-level information from the publicly Chicago Health Atlas with EMR data improved model fairness and unexpectedly revealed heterogeneous effects of neighborhood information across demographic groups. To further reduce model bias, FIG (Fairness through Intersectional Grouping) was developed as a model-agnostic multi-dimensional post-processing framework that systematically identifies domain-specific non-linear biased feature combinations and applies group-specific probability adjustments to reduce disparities across intersectional subgroups. The framework substantially reduced intersectional bias while maintaining predictive model performance, outperforming single-attribute fairness methods. Together, this work advances both the technical foundations of fair machine learning and its practical application in clinical prediction, demonstrating how integrating neighborhood context and addressing intersectional bias can improve model performance across diverse patient populations in healthcare settings.","abstract_html":"Machine learning models for predicting perinatal depression (PND) from electronic medical records (EMRs) show promise for early intervention, yet significant performance disparities (model bias) across sociodemographic groups limit their clinical utility for diverse populations. Current EMR-based approaches remain unexplored for low-income minority populations and fail to incorporate neighborhood-level information (e.g., pollution, access to healthy food, proximity to hospitals) that is well-known to affect human health. Further, existing computation methods to reduce model bias are designed to correct one attribute at a time (e.g., race, age, socioeconomic status), overlooking how combinations of patient attributes create unique patterns of prediction disparities (e.g., sex and race). This dissertation develops methodological approaches to enhance prediction fairness while maintaining clinical utility in predominantly low-income women of color at an urban academic medical center. Machine learning models were developed to predict early pregnancy depression while systematically assessing model bias using EMRs from individuals receiving obstetric care at UIHealth, where we mostly serve low-income women of color. Our models had worse performance for Black and Hispanic women despite their majority representation the data. Integration of neighborhood-level information from the publicly Chicago Health Atlas with EMR data improved model fairness and unexpectedly revealed heterogeneous effects of neighborhood information across demographic groups. To further reduce model bias, FIG (Fairness through Intersectional Grouping) was developed as a model-agnostic multi-dimensional post-processing framework that systematically identifies domain-specific non-linear biased feature combinations and applies group-specific probability adjustments to reduce disparities across intersectional subgroups. The framework substantially reduced intersectional bias while maintaining predictive model performance, outperforming single-attribute fairness methods. Together, this work advances both the technical foundations of fair machine learning and its practical application in clinical prediction, demonstrating how integrating neighborhood context and addressing intersectional bias can improve model performance across diverse patient populations in healthcare settings.","abstract_has_math":false,"creators":["Yongchao Huang (1449682)"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2026,"date_issued":"2026-05-01T00:00:00Z","date_published":"2026-05-01T00:00:00Z","updated_at":"2026-07-27T21:33:46Z","subjects":["Bioinformatics"],"languages":[],"rights":["In Copyright","Open Access after 2028-05-01"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.25417/uic.32994962.v1","outbound_label":"DOI","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Yongchao Huang (1449682)"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2026-05-01T00:00:00Z"]},{"key":"dc:relation","label":"Dc Relation","values":["https://figshare.com/articles/thesis/Advancing_Precision_in_Prenatal_Depression_Prediction_Leveraging_EMRs_and_Enhancing_ML-Model_Fairness/32994962"]},{"key":"dc:type","label":"Dc Type","values":["Text","Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Bioinformatics"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["In Copyright","Open Access after 2028-05-01"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["10.25417/uic.32994962.v1"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Machine learning models for predicting perinatal depression (PND) from electronic medical records (EMRs) show promise for early intervention, yet significant performance disparities (model bias) across sociodemographic groups limit their clinical utility for diverse populations. Current EMR-based approaches remain unexplored for low-income minority populations and fail to incorporate neighborhood-level information (e.g., pollution, access to healthy food, proximity to hospitals) that is well-known to affect human health. Further, existing computation methods to reduce model bias are designed to correct one attribute at a time (e.g., race, age, socioeconomic status), overlooking how combinations of patient attributes create unique patterns of prediction disparities (e.g., sex and race). This dissertation develops methodological approaches to enhance prediction fairness while maintaining clinical utility in predominantly low-income women of color at an urban academic medical center. Machine learning models were developed to predict early pregnancy depression while systematically assessing model bias using EMRs from individuals receiving obstetric care at UIHealth, where we mostly serve low-income women of color. Our models had worse performance for Black and Hispanic women despite their majority representation the data. Integration of neighborhood-level information from the publicly Chicago Health Atlas with EMR data improved model fairness and unexpectedly revealed heterogeneous effects of neighborhood information across demographic groups. To further reduce model bias, FIG (Fairness through Intersectional Grouping) was developed as a model-agnostic multi-dimensional post-processing framework that systematically identifies domain-specific non-linear biased feature combinations and applies group-specific probability adjustments to reduce disparities across intersectional subgroups. The framework substantially reduced intersectional bias while maintaining predictive model performance, outperforming single-attribute fairness methods. Together, this work advances both the technical foundations of fair machine learning and its practical application in clinical prediction, demonstrating how integrating neighborhood context and addressing intersectional bias can improve model performance across diverse patient populations in healthcare settings."]},{"key":"dc:title","label":"Title","values":["Advancing Precision in Prenatal Depression Prediction: Leveraging EMRs and Enhancing ML-Model Fairness"]}]}],"canonical_facts":{"dc:creator":["Yongchao Huang (1449682)"],"dc:date":["2026-05-01T00:00:00Z"],"dc:description":["Machine learning models for predicting perinatal depression (PND) from electronic medical records (EMRs) show promise for early intervention, yet significant performance disparities (model bias) across sociodemographic groups limit their clinical utility for diverse populations. Current EMR-based approaches remain unexplored for low-income minority populations and fail to incorporate neighborhood-level information (e.g., pollution, access to healthy food, proximity to hospitals) that is well-known to affect human health. Further, existing computation methods to reduce model bias are designed to correct one attribute at a time (e.g., race, age, socioeconomic status), overlooking how combinations of patient attributes create unique patterns of prediction disparities (e.g., sex and race). This dissertation develops methodological approaches to enhance prediction fairness while maintaining clinical utility in predominantly low-income women of color at an urban academic medical center. Machine learning models were developed to predict early pregnancy depression while systematically assessing model bias using EMRs from individuals receiving obstetric care at UIHealth, where we mostly serve low-income women of color. Our models had worse performance for Black and Hispanic women despite their majority representation the data. Integration of neighborhood-level information from the publicly Chicago Health Atlas with EMR data improved model fairness and unexpectedly revealed heterogeneous effects of neighborhood information across demographic groups. To further reduce model bias, FIG (Fairness through Intersectional Grouping) was developed as a model-agnostic multi-dimensional post-processing framework that systematically identifies domain-specific non-linear biased feature combinations and applies group-specific probability adjustments to reduce disparities across intersectional subgroups. The framework substantially reduced intersectional bias while maintaining predictive model performance, outperforming single-attribute fairness methods. Together, this work advances both the technical foundations of fair machine learning and its practical application in clinical prediction, demonstrating how integrating neighborhood context and addressing intersectional bias can improve model performance across diverse patient populations in healthcare settings."],"dc:identifier":["10.25417/uic.32994962.v1"],"dc:relation":["https://figshare.com/articles/thesis/Advancing_Precision_in_Prenatal_Depression_Prediction_Leveraging_EMRs_and_Enhancing_ML-Model_Fairness/32994962"],"dc:rights":["In Copyright","Open Access after 2028-05-01"],"dc:subject":["Bioinformatics"],"dc:title":["Advancing Precision in Prenatal Depression Prediction: Leveraging EMRs and Enhancing ML-Model Fairness"],"dc:type":["Text","Thesis"]},"updated_at":"2026-07-27T21:33:46Z"}