Back to results

Cornell University

Optimal and Safe Semi-supervised Estimation and Inference for High-dimensional Linear Regression

Abstract

dc:description.abstract

There are many scenarios such as the electronic health records where the outcome is much more difficult to collect than the covariates. We refer data consisting of both covariates and the corresponding outcomes as labeled data, and data with only covariates as unlabeled data. Semi-supervised learning combines both labeled and unlabeled data to improve a model using only labeled data and can be useful in these scenarios. In this work, we consider the linear regression problem with a semi-supervised learning data structure under high dimensionality. Our goal is to investigate when and how the unlabeled data can be exploited to improve the estimation and inference of the regression parameters in linear models, especially in light of the fact that linear models may be misspecified for real-world data. This work addresses the following two questions: (1) can we use the labeled data as well as the unlabeled data to construct a semi-supervised estimator such that its convergence rate is faster than the supervised estimators? (2) can we construct confidence intervals or hypothesis tests that are guaranteed to be more efficient or powerful than those with the supervised estimators? For the first question, we establish the minimax lower bound for the linear regression coefficients estimation in the semi-supervised setting and show that this lower bound cannot be achieved by supervised estimators using the labeled data only. We close this gap by proposing an optimal semi-supervised estimator with estimates for the conditional mean function as pseudo-labels for unlabeled data, which attains the lower bound provided that the imputation for the conditional mean function is consistent with a proper rate. To tackle the problem that without any model assumptions for the conditional mean, we cannot tell whether the imputation is consistent or not, we further propose a safe semi-supervised estimator. We view it safe, because this estimator is always at least as good as the supervised estimators regardless of the quality of the imputation. We also extend our idea to the aggregation of multiple semi-supervised estimators caused by different misspecifications of the conditional mean function. This allows us to borrow the predictive strength from a best imputation model to improve the supervised estimator. To answer the second question, based on the optimal semi-supervised estimator, we propose the efficient estimator for semi-supervised inference, which is fully efficient if the unknown conditional mean function is estimated consistently, but may not be more efficient than the supervised approach otherwise. We also further develop a safe inference procedure, which usually does not aim to provide fully efficient inference, but is guaranteed to be no worse than the supervised approach, no matter whether the linear model is correctly specified or the conditional mean function is consistently estimated. Extensive numerical simulations and real data analysis are conducted to illustrate our theoretical results.

Degree

thesis:*
Name thesis:degree_name
Ph. D., Statistics
Level thesis:degree_level
Doctor of Philosophy
Discipline thesis:degree_discipline
Statistics
Grantor
Cornell University
Year dc:date.issued
2023

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Deng, Siyi
Committee members dc:contributor.committeemember
  • Bunea, Florentina
  • Wells, Martin

Rights

Language dc:language.iso
en

Identifiers

dc:identifier.*
Dc Identifier Other
ProQuest Submission ID: 13671
ProQuest Publication ID: 30486400
OAI identifier oai:identifier
oai:ecommons.cornell.edu:1813/114019

Chain of custody

source
Harvested from
Cornell University
Base URL
ecommons.cornell.edu/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Deng, Siyi. Optimal and Safe Semi-supervised Estimation and Inference for High-dimensional Linear Regression. Doctor of Philosophy thesis, Cornell University, 2023. https://hdl.handle.net/1813/114019