{"id":{"repo_id":"cornell","oai_identifier":"oai:ecommons.cornell.edu:1813/114019"},"canonical_url":"https://search.dev.ndltd.org/etd/cornell/oai:ecommons.cornell.edu:1813/114019","repository":{"repo_id":"cornell","name":"Cornell University","base_url":"https://ecommons.cornell.edu/server/oai/request"},"display":{"title":"Optimal and Safe Semi-supervised Estimation and Inference for High-dimensional Linear Regression","abstract":"There are many scenarios such as the electronic health records where the outcome is much more difficult to collect than the covariates. We refer data consisting of both covariates and the corresponding outcomes as labeled data, and data with only covariates as unlabeled data. Semi-supervised learning combines both labeled and unlabeled data to improve a model using only labeled data and can be useful in these scenarios. In this work, we consider the linear regression problem with a semi-supervised learning data structure under high dimensionality. Our goal is to investigate when and how the unlabeled data can be exploited to improve the estimation and inference of the regression parameters in linear models, especially in light of the fact that linear models may be misspecified for real-world data. This work addresses the following two questions: (1) can we use the labeled data as well as the unlabeled data to construct a semi-supervised estimator such that its convergence rate is faster than the supervised estimators? (2) can we construct confidence intervals or hypothesis tests that are guaranteed to be more efficient or powerful than those with the supervised estimators? For the first question, we establish the minimax lower bound for the linear regression coefficients estimation in the semi-supervised setting and show that this lower bound cannot be achieved by supervised estimators using the labeled data only. We close this gap by proposing an optimal semi-supervised estimator with estimates for the conditional mean function as pseudo-labels for unlabeled data, which attains the lower bound provided that the imputation for the conditional mean function is consistent with a proper rate. To tackle the problem that without any model assumptions for the conditional mean, we cannot tell whether the imputation is consistent or not, we further propose a safe semi-supervised estimator. We view it safe, because this estimator is always at least as good as the supervised estimators regardless of the quality of the imputation. We also extend our idea to the aggregation of multiple semi-supervised estimators caused by different misspecifications of the conditional mean function. This allows us to borrow the predictive strength from a best imputation model to improve the supervised estimator. To answer the second question, based on the optimal semi-supervised estimator, we propose the efficient estimator for semi-supervised inference, which is fully efficient if the unknown conditional mean function is estimated consistently, but may not be more efficient than the supervised approach otherwise. We also further develop a safe inference procedure, which usually does not aim to provide fully efficient inference, but is guaranteed to be no worse than the supervised approach, no matter whether the linear model is correctly specified or the conditional mean function is consistently estimated. Extensive numerical simulations and real data analysis are conducted to illustrate our theoretical results.","abstract_html":"There are many scenarios such as the electronic health records where the outcome is much more difficult to collect than the covariates. We refer data consisting of both covariates and the corresponding outcomes as labeled data, and data with only covariates as unlabeled data. Semi-supervised learning combines both labeled and unlabeled data to improve a model using only labeled data and can be useful in these scenarios. In this work, we consider the linear regression problem with a semi-supervised learning data structure under high dimensionality. Our goal is to investigate when and how the unlabeled data can be exploited to improve the estimation and inference of the regression parameters in linear models, especially in light of the fact that linear models may be misspecified for real-world data. This work addresses the following two questions: (1) can we use the labeled data as well as the unlabeled data to construct a semi-supervised estimator such that its convergence rate is faster than the supervised estimators? (2) can we construct confidence intervals or hypothesis tests that are guaranteed to be more efficient or powerful than those with the supervised estimators? For the first question, we establish the minimax lower bound for the linear regression coefficients estimation in the semi-supervised setting and show that this lower bound cannot be achieved by supervised estimators using the labeled data only. We close this gap by proposing an optimal semi-supervised estimator with estimates for the conditional mean function as pseudo-labels for unlabeled data, which attains the lower bound provided that the imputation for the conditional mean function is consistent with a proper rate. To tackle the problem that without any model assumptions for the conditional mean, we cannot tell whether the imputation is consistent or not, we further propose a safe semi-supervised estimator. We view it safe, because this estimator is always at least as good as the supervised estimators regardless of the quality of the imputation. We also extend our idea to the aggregation of multiple semi-supervised estimators caused by different misspecifications of the conditional mean function. This allows us to borrow the predictive strength from a best imputation model to improve the supervised estimator. To answer the second question, based on the optimal semi-supervised estimator, we propose the efficient estimator for semi-supervised inference, which is fully efficient if the unknown conditional mean function is estimated consistently, but may not be more efficient than the supervised approach otherwise. We also further develop a safe inference procedure, which usually does not aim to provide fully efficient inference, but is guaranteed to be no worse than the supervised approach, no matter whether the linear model is correctly specified or the conditional mean function is consistently estimated. Extensive numerical simulations and real data analysis are conducted to illustrate our theoretical results.","abstract_has_math":false,"creators":["Deng, Siyi"],"institution":"Cornell University","degree_name":"Ph. D., Statistics","degree_level":"Doctor of Philosophy","degree_discipline":"Statistics","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":["Bunea, Florentina","Wells, Martin"],"year":2023,"date_issued":"2023-05","date_published":"2023-05","updated_at":"2026-07-24T01:48:56Z","subjects":[],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.7298/znzx-ks27"],"render_values":[{"text":"https://doi.org/10.7298/znzx-ks27","href":"https://doi.org/10.7298/znzx-ks27","code":true}]},{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["ProQuest Submission ID: 13671","ProQuest Publication ID: 30486400"],"render_values":[{"text":"ProQuest Submission ID: 13671","href":null,"code":true},{"text":"ProQuest Publication ID: 30486400","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/1813/114019","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Bunea, Florentina","Wells, Martin"]},{"key":"dc:creator","label":"Author","values":["Deng, Siyi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2024-01-31T21:18:48Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2024-01-31T21:18:48Z"]},{"key":"dc:date.issued","label":"Date","values":["2023-05"]},{"key":"dc:type","label":"Dc Type","values":["dissertation or thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctor of Philosophy"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph. D., Statistics"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Cornell University"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.7298/znzx-ks27"]},{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["ProQuest Submission ID: 13671","ProQuest Publication ID: 30486400"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1813/114019"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["There are many scenarios such as the electronic health records where the outcome is much more difficult to collect than the covariates. We refer data consisting of both covariates and the corresponding outcomes as labeled data, and data with only covariates as unlabeled data. Semi-supervised learning combines both labeled and unlabeled data to improve a model using only labeled data and can be useful in these scenarios. In this work, we consider the linear regression problem with a semi-supervised learning data structure under high dimensionality. Our goal is to investigate when and how the unlabeled data can be exploited to improve the estimation and inference of the regression parameters in linear models, especially in light of the fact that linear models may be misspecified for real-world data. This work addresses the following two questions: (1) can we use the labeled data as well as the unlabeled data to construct a semi-supervised estimator such that its convergence rate is faster than the supervised estimators? (2) can we construct confidence intervals or hypothesis tests that are guaranteed to be more efficient or powerful than those with the supervised estimators? For the first question, we establish the minimax lower bound for the linear regression coefficients estimation in the semi-supervised setting and show that this lower bound cannot be achieved by supervised estimators using the labeled data only. We close this gap by proposing an optimal semi-supervised estimator with estimates for the conditional mean function as pseudo-labels for unlabeled data, which attains the lower bound provided that the imputation for the conditional mean function is consistent with a proper rate. To tackle the problem that without any model assumptions for the conditional mean, we cannot tell whether the imputation is consistent or not, we further propose a safe semi-supervised estimator. We view it safe, because this estimator is always at least as good as the supervised estimators regardless of the quality of the imputation. We also extend our idea to the aggregation of multiple semi-supervised estimators caused by different misspecifications of the conditional mean function. This allows us to borrow the predictive strength from a best imputation model to improve the supervised estimator. To answer the second question, based on the optimal semi-supervised estimator, we propose the efficient estimator for semi-supervised inference, which is fully efficient if the unknown conditional mean function is estimated consistently, but may not be more efficient than the supervised approach otherwise. We also further develop a safe inference procedure, which usually does not aim to provide fully efficient inference, but is guaranteed to be no worse than the supervised approach, no matter whether the linear model is correctly specified or the conditional mean function is consistently estimated. Extensive numerical simulations and real data analysis are conducted to illustrate our theoretical results."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Optimal and Safe Semi-supervised Estimation and Inference for High-dimensional Linear Regression"]}]}],"canonical_facts":{"dc:contributor.committeemember":["Bunea, Florentina","Wells, Martin"],"dc:creator":["Deng, Siyi"],"dc:date.accessioned":["2024-01-31T21:18:48Z"],"dc:date.available":["2024-01-31T21:18:48Z"],"dc:date.issued":["2023-05"],"dc:description.abstract":["There are many scenarios such as the electronic health records where the outcome is much more difficult to collect than the covariates. We refer data consisting of both covariates and the corresponding outcomes as labeled data, and data with only covariates as unlabeled data. Semi-supervised learning combines both labeled and unlabeled data to improve a model using only labeled data and can be useful in these scenarios. In this work, we consider the linear regression problem with a semi-supervised learning data structure under high dimensionality. Our goal is to investigate when and how the unlabeled data can be exploited to improve the estimation and inference of the regression parameters in linear models, especially in light of the fact that linear models may be misspecified for real-world data. This work addresses the following two questions: (1) can we use the labeled data as well as the unlabeled data to construct a semi-supervised estimator such that its convergence rate is faster than the supervised estimators? (2) can we construct confidence intervals or hypothesis tests that are guaranteed to be more efficient or powerful than those with the supervised estimators? For the first question, we establish the minimax lower bound for the linear regression coefficients estimation in the semi-supervised setting and show that this lower bound cannot be achieved by supervised estimators using the labeled data only. We close this gap by proposing an optimal semi-supervised estimator with estimates for the conditional mean function as pseudo-labels for unlabeled data, which attains the lower bound provided that the imputation for the conditional mean function is consistent with a proper rate. To tackle the problem that without any model assumptions for the conditional mean, we cannot tell whether the imputation is consistent or not, we further propose a safe semi-supervised estimator. We view it safe, because this estimator is always at least as good as the supervised estimators regardless of the quality of the imputation. We also extend our idea to the aggregation of multiple semi-supervised estimators caused by different misspecifications of the conditional mean function. This allows us to borrow the predictive strength from a best imputation model to improve the supervised estimator. To answer the second question, based on the optimal semi-supervised estimator, we propose the efficient estimator for semi-supervised inference, which is fully efficient if the unknown conditional mean function is estimated consistently, but may not be more efficient than the supervised approach otherwise. We also further develop a safe inference procedure, which usually does not aim to provide fully efficient inference, but is guaranteed to be no worse than the supervised approach, no matter whether the linear model is correctly specified or the conditional mean function is consistently estimated. Extensive numerical simulations and real data analysis are conducted to illustrate our theoretical results."],"dc:format.mimetype":["application/pdf"],"dc:identifier.doi":["https://doi.org/10.7298/znzx-ks27"],"dc:identifier.other":["ProQuest Submission ID: 13671","ProQuest Publication ID: 30486400"],"dc:identifier.uri":["https://hdl.handle.net/1813/114019"],"dc:language.iso":["en"],"dc:title":["Optimal and Safe Semi-supervised Estimation and Inference for High-dimensional Linear Regression"],"dc:type":["dissertation or thesis"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["Doctor of Philosophy"],"thesis:degree_name":["Ph. D., Statistics"],"thesis:institution_name":["Cornell University"]},"updated_at":"2026-07-24T01:48:56Z"}