{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/127347"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/127347","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Subpopulation selection and debiased estimation for causal inference and predictive model evaluation","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-12-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2026-12-01","abstract_has_math":false,"creators":["Liu, Yuxuan"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Statistics","degree_department":null,"school":null,"contributors":["Simpson, Douglas","Li, Xinran","Zhu, Ruoqing","Culpepper, Steven Andrew"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-11-25","date_published":"2024-11-25","updated_at":"2026-07-22T22:25:03Z","subjects":["Causal Inference","Debiased Estimation","Classification Evaluation"],"languages":["en","eng"],"rights":["© 2024 Yuxuan Liu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/127347","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Simpson, Douglas","Li, Xinran","Zhu, Ruoqing","Culpepper, Steven Andrew"]},{"key":"dc:creator","label":"Author","values":["Liu, Yuxuan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-11-25","2024-12"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Causal Inference","Debiased Estimation","Classification Evaluation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["© 2024 Yuxuan Liu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/127347"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-12-01","The student, Yuxuan Liu, accepted the attached license on 2024-11-14 at 15:18.","The student, Yuxuan Liu, submitted this Dissertation for approval on 2024-11-14 at 15:44.","This Dissertation was approved for publication on 2024-11-25 at 11:59.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21320 on 2025-03-28 at 14:43:05","Bias in statistical and machine learning models refers to systematic errors that can skew results, leading to inaccurate conclusions and predictions. It can arise from various sources, including data collection methods, model selection, and underlying assumptions. Debiasing techniques aim to mitigate these errors and improve the reliability and validity of models. The goal of debiasing is to enhance model performance by reducing systematic errors, thus providing more accurate and trustworthy results for decision-making and inference. Common strategies for debiasing include cross-validation and bootstrap. Other methods involve correcting for biases introduced by extreme values, imbalanced datasets, and confounding variables. In this thesis, two chapters focus on advancing estimator debias methodologies in causal inference and classification, particularly in the realms of average treatment effect estimation and classification model evaluation. The first chapter introduces a method for interpretable weighted average treatment effect (WATE) estimation under possible violation of overlapping assumption. By selecting subpopula- tions through optimizing a loss function via integer linear programming, the approach enhances the precision and reliability of treatment effect estimates, especially in the presence of extreme propensity scores. Simulation studies show that the proposed WATE estimators outperform classic IPW estimators in terms of bias reduction by a significant magnitude. In the second chapter, classification model evaluation is examined, particularly being focused on cross-validation and bootstrap methods. Previously, our case study demonstrates the enhanced prediction of preterm birth that incorporates two nested cross-validated machine learning classifi- cation models. As a beginning, this chapter delves into the theoretical aspects of cross-validation for linear regression, confirming that sample mean square error is an unbiased consistent estimator of true MSE. After so, this chapter explores the AUC metric for evaluating various classifiers performances and addresses the limitations of the DeLong test, proposing alternatives like bootstrap resampling. Finally, this chapter highlights the issue of bias in nested model AUC difference estimation and demonstrates that sample splitting involvement provides unbiased, asymptotically normal estimators for cross-validated generalized linear model (GLM) and linear discriminant analysis (LDA)."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Subpopulation selection and debiased estimation for causal inference and predictive model evaluation"]}]}],"canonical_facts":{"dc:contributor":["Simpson, Douglas","Li, Xinran","Zhu, Ruoqing","Culpepper, Steven Andrew"],"dc:creator":["Liu, Yuxuan"],"dc:date":["2024-11-25","2024-12"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-12-01","The student, Yuxuan Liu, accepted the attached license on 2024-11-14 at 15:18.","The student, Yuxuan Liu, submitted this Dissertation for approval on 2024-11-14 at 15:44.","This Dissertation was approved for publication on 2024-11-25 at 11:59.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21320 on 2025-03-28 at 14:43:05","Bias in statistical and machine learning models refers to systematic errors that can skew results, leading to inaccurate conclusions and predictions. It can arise from various sources, including data collection methods, model selection, and underlying assumptions. Debiasing techniques aim to mitigate these errors and improve the reliability and validity of models. The goal of debiasing is to enhance model performance by reducing systematic errors, thus providing more accurate and trustworthy results for decision-making and inference. Common strategies for debiasing include cross-validation and bootstrap. Other methods involve correcting for biases introduced by extreme values, imbalanced datasets, and confounding variables. In this thesis, two chapters focus on advancing estimator debias methodologies in causal inference and classification, particularly in the realms of average treatment effect estimation and classification model evaluation. The first chapter introduces a method for interpretable weighted average treatment effect (WATE) estimation under possible violation of overlapping assumption. By selecting subpopula- tions through optimizing a loss function via integer linear programming, the approach enhances the precision and reliability of treatment effect estimates, especially in the presence of extreme propensity scores. Simulation studies show that the proposed WATE estimators outperform classic IPW estimators in terms of bias reduction by a significant magnitude. In the second chapter, classification model evaluation is examined, particularly being focused on cross-validation and bootstrap methods. Previously, our case study demonstrates the enhanced prediction of preterm birth that incorporates two nested cross-validated machine learning classifi- cation models. As a beginning, this chapter delves into the theoretical aspects of cross-validation for linear regression, confirming that sample mean square error is an unbiased consistent estimator of true MSE. After so, this chapter explores the AUC metric for evaluating various classifiers performances and addresses the limitations of the DeLong test, proposing alternatives like bootstrap resampling. Finally, this chapter highlights the issue of bias in nested model AUC difference estimation and demonstrates that sample splitting involvement provides unbiased, asymptotically normal estimators for cross-validated generalized linear model (GLM) and linear discriminant analysis (LDA)."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/127347"],"dc:language":["en","eng"],"dc:rights":["© 2024 Yuxuan Liu"],"dc:subject":["Causal Inference","Debiased Estimation","Classification Evaluation"],"dc:title":["Subpopulation selection and debiased estimation for causal inference and predictive model evaluation"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:03Z"}