{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/137106"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/137106","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Methodologies for Systematic Evaluation and Targeted Mitigation of Deficiencies in Critical Machine Learning Models","abstract":"Despite the growing use of machine learning in healthcare, critical challenges remain unaddressed, models often fail to respond appropriately to life-threatening conditions, exhibit poor generalizability in real-world clinical settings, and show unequal performance across patient subgroups. These limitations compromise the reliability, safety, and equity of AI-driven decision-making, especially in high-stakes environments like intensive care. In this work, we outline a comprehensive evaluation and mitigation strategy to address both responsiveness and fairness shortcomings.. We develop testing approaches to systematically assess models' ability to respond to serious medical emergencies. Using generated test cases, we found that statistical machine-learning models trained solely from patient data are grossly insufficient and have many dangerous blind spots. Specifically, we identified serious deficiencies in the models' responsiveness, i.e., the inability to recognize severely impaired medical conditions or rapidly deteriorating health. For in-hospital mortality prediction, the models tested using our synthesized cases fail to recognize 66% of the test cases involving injuries. In some instances, the models fail to generate adequate mortality risk scores for all test cases. We also applied our testing methods to assess the responsiveness of 5-year breast and lung cancer prediction models and identified similar kinds of deficiencies. To address the low responsiveness of machine learning models to critical health conditions, we integrated domain knowledge into the modeling framework using two complementary strategies: (i) a custom loss function that penalizes violations of medical constraints, and (ii) a rule-based decision tree derived from clinical knowledge, aggregated with a data-driven model. The resulting knowledge-guided models demonstrated notable improvements in performance, particularly under critical scenarios. For instance, recall improved by 7% on the full glucose test set and by 27% for critically high glucose cases, achieving 94–99% accuracy in detecting patients with severely abnormal glucose levels. Similar trends were observed for other vital signs. Moreover, the decision tree-based hybrid model enhanced early sepsis detection accuracy by 4%, underscoring the benefit of combining clinical knowledge with statistical learning for high-stakes medical applications. In addition, we address a bias problem we identified in models predicting type 2 diabetes, which disproportionately impacts younger adults, a growing segment of diabetes patients. In this research, we identify this deficiency in traditional machine learning models and propose an algorithm to mitigate the bias towards the young population when predicting diabetes. Deviating from the traditional concept of one-model-fits-all, we train customized machine-learning models for each age group. Our proposed solution consistently improves recall of diabetes class by 26% to 40% in the young age group (30-44). Moreover, our technique outperforms 7 commonly used whole-group sampling techniques such as random oversampling, SMOTE, and AdaSyns techniques by at least 36% in terms of diabetes recall in the young age group.","abstract_html":"Despite the growing use of machine learning in healthcare, critical challenges remain unaddressed, models often fail to respond appropriately to life-threatening conditions, exhibit poor generalizability in real-world clinical settings, and show unequal performance across patient subgroups. These limitations compromise the reliability, safety, and equity of AI-driven decision-making, especially in high-stakes environments like intensive care. In this work, we outline a comprehensive evaluation and mitigation strategy to address both responsiveness and fairness shortcomings.. We develop testing approaches to systematically assess models&#x27; ability to respond to serious medical emergencies. Using generated test cases, we found that statistical machine-learning models trained solely from patient data are grossly insufficient and have many dangerous blind spots. Specifically, we identified serious deficiencies in the models&#x27; responsiveness, i.e., the inability to recognize severely impaired medical conditions or rapidly deteriorating health. For in-hospital mortality prediction, the models tested using our synthesized cases fail to recognize 66% of the test cases involving injuries. In some instances, the models fail to generate adequate mortality risk scores for all test cases. We also applied our testing methods to assess the responsiveness of 5-year breast and lung cancer prediction models and identified similar kinds of deficiencies. To address the low responsiveness of machine learning models to critical health conditions, we integrated domain knowledge into the modeling framework using two complementary strategies: (i) a custom loss function that penalizes violations of medical constraints, and (ii) a rule-based decision tree derived from clinical knowledge, aggregated with a data-driven model. The resulting knowledge-guided models demonstrated notable improvements in performance, particularly under critical scenarios. For instance, recall improved by 7% on the full glucose test set and by 27% for critically high glucose cases, achieving 94–99% accuracy in detecting patients with severely abnormal glucose levels. Similar trends were observed for other vital signs. Moreover, the decision tree-based hybrid model enhanced early sepsis detection accuracy by 4%, underscoring the benefit of combining clinical knowledge with statistical learning for high-stakes medical applications. In addition, we address a bias problem we identified in models predicting type 2 diabetes, which disproportionately impacts younger adults, a growing segment of diabetes patients. In this research, we identify this deficiency in traditional machine learning models and propose an algorithm to mitigate the bias towards the young population when predicting diabetes. Deviating from the traditional concept of one-model-fits-all, we train customized machine-learning models for each age group. Our proposed solution consistently improves recall of diabetes class by 26% to 40% in the young age group (30-44). Moreover, our technique outperforms 7 commonly used whole-group sampling techniques such as random oversampling, SMOTE, and AdaSyns techniques by at least 36% in terms of diabetes recall in the young age group.","abstract_has_math":false,"creators":["Pias, Tanmoy Sarkar"],"institution":"Virginia Tech","degree_name":"Doctor of Philosophy","degree_level":"doctoral","degree_discipline":"Computer Science & Applications","degree_department":"Computer Science and#38; Applications","school":null,"contributors":[],"advisors":[],"committee_chairs":["Yao, Danfeng"],"committee_members":["Chiu, Pearl Huh","Joshi, Shalmali Dilip","Murali, T. M.","Lourentzou, Ismini"],"year":2025,"date_issued":"2025-08-07","date_published":"2025-08-07","updated_at":"2026-07-22T22:19:04Z","subjects":["AI Trustworthiness","Responsiveness","Knowledge guided ML","Custom Loss","Healthcare"],"languages":["en"],"rights":["Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International"],"rights_urls":["http://creativecommons.org/licenses/by-nc-nd/4.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:44496"],"render_values":[{"text":"vt_gsexam:44496","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10919/137106","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Yao, Danfeng"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Chiu, Pearl Huh","Joshi, Shalmali Dilip","Murali, T. M.","Lourentzou, Ismini"]},{"key":"dc:contributor.department","label":"Department","values":["Computer Science and#38; Applications"]},{"key":"dc:creator","label":"Author","values":["Pias, Tanmoy Sarkar"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-08-08T08:00:51Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-08-08T08:00:51Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-08-07"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science & Applications"]},{"key":"thesis:degree_level","label":"Degree Level","values":["doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["AI Trustworthiness","Responsiveness","Knowledge guided ML","Custom Loss","Healthcare"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://creativecommons.org/licenses/by-nc-nd/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:44496"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/137106"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Despite the growing use of machine learning in healthcare, critical challenges remain unaddressed, models often fail to respond appropriately to life-threatening conditions, exhibit poor generalizability in real-world clinical settings, and show unequal performance across patient subgroups. These limitations compromise the reliability, safety, and equity of AI-driven decision-making, especially in high-stakes environments like intensive care. In this work, we outline a comprehensive evaluation and mitigation strategy to address both responsiveness and fairness shortcomings.. We develop testing approaches to systematically assess models' ability to respond to serious medical emergencies. Using generated test cases, we found that statistical machine-learning models trained solely from patient data are grossly insufficient and have many dangerous blind spots. Specifically, we identified serious deficiencies in the models' responsiveness, i.e., the inability to recognize severely impaired medical conditions or rapidly deteriorating health. For in-hospital mortality prediction, the models tested using our synthesized cases fail to recognize 66% of the test cases involving injuries. In some instances, the models fail to generate adequate mortality risk scores for all test cases. We also applied our testing methods to assess the responsiveness of 5-year breast and lung cancer prediction models and identified similar kinds of deficiencies. To address the low responsiveness of machine learning models to critical health conditions, we integrated domain knowledge into the modeling framework using two complementary strategies: (i) a custom loss function that penalizes violations of medical constraints, and (ii) a rule-based decision tree derived from clinical knowledge, aggregated with a data-driven model. The resulting knowledge-guided models demonstrated notable improvements in performance, particularly under critical scenarios. For instance, recall improved by 7% on the full glucose test set and by 27% for critically high glucose cases, achieving 94–99% accuracy in detecting patients with severely abnormal glucose levels. Similar trends were observed for other vital signs. Moreover, the decision tree-based hybrid model enhanced early sepsis detection accuracy by 4%, underscoring the benefit of combining clinical knowledge with statistical learning for high-stakes medical applications. In addition, we address a bias problem we identified in models predicting type 2 diabetes, which disproportionately impacts younger adults, a growing segment of diabetes patients. In this research, we identify this deficiency in traditional machine learning models and propose an algorithm to mitigate the bias towards the young population when predicting diabetes. Deviating from the traditional concept of one-model-fits-all, we train customized machine-learning models for each age group. Our proposed solution consistently improves recall of diabetes class by 26% to 40% in the young age group (30-44). Moreover, our technique outperforms 7 commonly used whole-group sampling techniques such as random oversampling, SMOTE, and AdaSyns techniques by at least 36% in terms of diabetes recall in the young age group."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["Despite the growing use of machine learning in healthcare, critical challenges remain unaddressed, models often fail to respond appropriately to life-threatening conditions, exhibit poor generalizability in real-world clinical settings, and show unequal performance across patient subgroups. These limitations compromise the reliability, safety, and equity of AI-driven decision-making, especially in high-stakes environments like intensive care. In this work, we outline a comprehensive evaluation and mitigation strategy to address both responsiveness and fairness shortcomings. In this research, we develop new methods to test ML models under critical health scenarios. Our findings reveal that current models, trained solely on patient data, have significant blind spots; many fail to recognize severe conditions, accurately predict mortality, or assess injury-related risks. For instance, in our tests, the models missed 66% of injury-related cases and often provided inadequate risk scores for patients who were actually at high risk. We observed similar limitations in models predicting long-term survival for breast and lung cancer, highlighting widespread responsiveness issues in current healthcare ML tools. To address the low responsiveness of machine learning models to critical health conditions, we integrated domain knowledge into the modeling framework using two complementary strategies: (i) a custom loss function that penalizes violations of medical constraints, and (ii) a rule-based decision tree derived from clinical knowledge, aggregated with a data-driven model. The resulting knowledge-guided models demonstrated notable improvements in performance, particularly under critical scenarios. For instance, recall improved by 7% on the full glucose test set and by 27% for critically high glucose cases, achieving 94–99% accuracy in detecting patients with severely abnormal glucose levels. Similar trends were observed for other vital signs. Moreover, the decision tree-based hybrid model enhanced early sepsis detection accuracy by 4%, underscoring the benefit of combining clinical knowledge with statistical learning for high-stakes medical applications. In addition, we address a bias problem we identified in models predicting type 2 diabetes, which disproportionately impacts younger adults, a growing segment of diabetes patients. Many traditional ML models exhibit \"digital ageism,\" or a tendency to overlook diabetes risk in younger populations. To counteract this, we designed age-specific models that improved detection accuracy by 26-40% for younger adults (ages 30-44) compared to conventional methods. Our approach also outperformed common techniques by at least 36% in recall for young adults, providing a more equitable solution for diabetes risk prediction."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Doctor of Philosophy"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["Methodologies for Systematic Evaluation and Targeted Mitigation of Deficiencies in Critical Machine Learning Models"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Yao, Danfeng"],"dc:contributor.committeemember":["Chiu, Pearl Huh","Joshi, Shalmali Dilip","Murali, T. M.","Lourentzou, Ismini"],"dc:contributor.department":["Computer Science and#38; Applications"],"dc:creator":["Pias, Tanmoy Sarkar"],"dc:date.accessioned":["2025-08-08T08:00:51Z"],"dc:date.available":["2025-08-08T08:00:51Z"],"dc:date.issued":["2025-08-07"],"dc:description.abstract":["Despite the growing use of machine learning in healthcare, critical challenges remain unaddressed, models often fail to respond appropriately to life-threatening conditions, exhibit poor generalizability in real-world clinical settings, and show unequal performance across patient subgroups. These limitations compromise the reliability, safety, and equity of AI-driven decision-making, especially in high-stakes environments like intensive care. In this work, we outline a comprehensive evaluation and mitigation strategy to address both responsiveness and fairness shortcomings.. We develop testing approaches to systematically assess models' ability to respond to serious medical emergencies. Using generated test cases, we found that statistical machine-learning models trained solely from patient data are grossly insufficient and have many dangerous blind spots. Specifically, we identified serious deficiencies in the models' responsiveness, i.e., the inability to recognize severely impaired medical conditions or rapidly deteriorating health. For in-hospital mortality prediction, the models tested using our synthesized cases fail to recognize 66% of the test cases involving injuries. In some instances, the models fail to generate adequate mortality risk scores for all test cases. We also applied our testing methods to assess the responsiveness of 5-year breast and lung cancer prediction models and identified similar kinds of deficiencies. To address the low responsiveness of machine learning models to critical health conditions, we integrated domain knowledge into the modeling framework using two complementary strategies: (i) a custom loss function that penalizes violations of medical constraints, and (ii) a rule-based decision tree derived from clinical knowledge, aggregated with a data-driven model. The resulting knowledge-guided models demonstrated notable improvements in performance, particularly under critical scenarios. For instance, recall improved by 7% on the full glucose test set and by 27% for critically high glucose cases, achieving 94–99% accuracy in detecting patients with severely abnormal glucose levels. Similar trends were observed for other vital signs. Moreover, the decision tree-based hybrid model enhanced early sepsis detection accuracy by 4%, underscoring the benefit of combining clinical knowledge with statistical learning for high-stakes medical applications. In addition, we address a bias problem we identified in models predicting type 2 diabetes, which disproportionately impacts younger adults, a growing segment of diabetes patients. In this research, we identify this deficiency in traditional machine learning models and propose an algorithm to mitigate the bias towards the young population when predicting diabetes. Deviating from the traditional concept of one-model-fits-all, we train customized machine-learning models for each age group. Our proposed solution consistently improves recall of diabetes class by 26% to 40% in the young age group (30-44). Moreover, our technique outperforms 7 commonly used whole-group sampling techniques such as random oversampling, SMOTE, and AdaSyns techniques by at least 36% in terms of diabetes recall in the young age group."],"dc:description.abstractgeneral":["Despite the growing use of machine learning in healthcare, critical challenges remain unaddressed, models often fail to respond appropriately to life-threatening conditions, exhibit poor generalizability in real-world clinical settings, and show unequal performance across patient subgroups. These limitations compromise the reliability, safety, and equity of AI-driven decision-making, especially in high-stakes environments like intensive care. In this work, we outline a comprehensive evaluation and mitigation strategy to address both responsiveness and fairness shortcomings. In this research, we develop new methods to test ML models under critical health scenarios. Our findings reveal that current models, trained solely on patient data, have significant blind spots; many fail to recognize severe conditions, accurately predict mortality, or assess injury-related risks. For instance, in our tests, the models missed 66% of injury-related cases and often provided inadequate risk scores for patients who were actually at high risk. We observed similar limitations in models predicting long-term survival for breast and lung cancer, highlighting widespread responsiveness issues in current healthcare ML tools. To address the low responsiveness of machine learning models to critical health conditions, we integrated domain knowledge into the modeling framework using two complementary strategies: (i) a custom loss function that penalizes violations of medical constraints, and (ii) a rule-based decision tree derived from clinical knowledge, aggregated with a data-driven model. The resulting knowledge-guided models demonstrated notable improvements in performance, particularly under critical scenarios. For instance, recall improved by 7% on the full glucose test set and by 27% for critically high glucose cases, achieving 94–99% accuracy in detecting patients with severely abnormal glucose levels. Similar trends were observed for other vital signs. Moreover, the decision tree-based hybrid model enhanced early sepsis detection accuracy by 4%, underscoring the benefit of combining clinical knowledge with statistical learning for high-stakes medical applications. In addition, we address a bias problem we identified in models predicting type 2 diabetes, which disproportionately impacts younger adults, a growing segment of diabetes patients. Many traditional ML models exhibit \"digital ageism,\" or a tendency to overlook diabetes risk in younger populations. To counteract this, we designed age-specific models that improved detection accuracy by 26-40% for younger adults (ages 30-44) compared to conventional methods. Our approach also outperformed common techniques by at least 36% in recall for young adults, providing a more equitable solution for diabetes risk prediction."],"dc:description.degree":["Doctor of Philosophy"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:44496"],"dc:identifier.uri":["https://hdl.handle.net/10919/137106"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International"],"dc:rights.uri":["http://creativecommons.org/licenses/by-nc-nd/4.0/"],"dc:subject":["AI Trustworthiness","Responsiveness","Knowledge guided ML","Custom Loss","Healthcare"],"dc:title":["Methodologies for Systematic Evaluation and Targeted Mitigation of Deficiencies in Critical Machine Learning Models"],"dc:type":["Dissertation"],"thesis:degree_discipline":["Computer Science & Applications"],"thesis:degree_level":["doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:19:04Z"}