{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/117562"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/117562","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Analysis based on incomplete data","abstract":"Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2024-12-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;Closed Access&#x27;, the embargo will last until 2024-12-01","abstract_has_math":false,"creators":["Qu, Tianyi"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Statistics","degree_department":null,"school":null,"contributors":["Li, Bo","Li, Xinran","Shao, Xiaofeng","Wang, Shulei"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-12","date_published":"2022-12","updated_at":"2026-07-22T22:24:56Z","subjects":["Likelihood","Matrix Factorization","Missing Value","Spatiotemporal Data","Randomization-based Inference","Potential Outcomes","M-estimation","Generalized Linear Models","Treatment Effect Heterogeneity","Model-assisted Approach","Bandit Problems","Adaptive Allocation","Metropolis-hasting Optimization"],"languages":["en","eng"],"rights":["Copyright 2022 Tianyi Qu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/117562","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Li, Bo","Li, Xinran","Shao, Xiaofeng","Wang, Shulei"]},{"key":"dc:creator","label":"Author","values":["Qu, Tianyi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-12","2022-11-29"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Likelihood","Matrix Factorization","Missing Value","Spatiotemporal Data","Randomization-based Inference","Potential Outcomes","M-estimation","Generalized Linear Models","Treatment Effect Heterogeneity","Model-assisted Approach","Bandit Problems","Adaptive Allocation","Metropolis-hasting Optimization"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Tianyi Qu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/117562"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2024-12-01","The student, Tianyi Qu, accepted the attached license on 2022-11-21 at 19:36.","The student, Tianyi Qu, submitted this Dissertation for approval on 2022-11-27 at 00:29.","This Dissertation was approved for publication on 2022-11-29 at 10:54.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18612 on 2023-04-12 at 11:35:07","In many situations, we may have to rely on partial information to do data analysis due to various reasons. For example, it may be that part of the data is not available, such as missing values and causal inference, or the analysis has to be drawn before the whole data is collected, such as bandit. In this thesis, we address three challenges that involve analysis with partial information. In the first chapter, we focus on prediction with unbalanced missing values. Public health data, such as HIV new diagnoses, are often left-censored due to confidentiality issues. Standard analysis approaches that assume censored values as missing at random often lead to biased estimates and inferior predictions. Motivated by the Philadelphia areal counts of HIV new diagnoses for which all values less than or equal to 5 are suppressed, we propose two methods to reduce the adverse influence of missingness on predictions and imputation of areal HIV new diagnoses. One is the likelihood-based method that integrates the missing mechanism into the likelihood function, and the other is a nonparametric algorithm for matrix factorization imputation. Numerical studies and the Philadelphia data analysis demonstrate that the two proposed methods can significantly improve prediction and imputation based on left-censored HIV data. We also compare the two methods on their robustness to model misspecification and find that both methods appear to be robust for prediction, while their performance for imputation depends on model specification. In the second chapter, we focus on the causal inference that for each individual, either treatment or control outcome can be observed. Randomized experiments have been the gold standard for drawing causal inferences. Conventional model-based analysis has been one of the most popular ways of analyzing treatment effects from randomized experiments, which is often carried through inference for certain model parameters. We provide a systematic investigation of model-based analysis, including the theory of M-estimation, under the randomization-based inference framework, avoiding any distributional assumptions on outcomes or covariates and utilizing only randomization as the ``reasoned basis''. We first show that the conventional model-based approach generally provides biased treatment effect estimation. We then study the model-imputed approach that uses the models mainly as a tool for imputing potential outcomes. Such an approach, although generally leading to biased estimation as well, can be valid for some special classes of models, e.g., the generalized linear model with canonical links. We finally recommend the model-assisted approach, which always provides consistent estimation and is robust to arbitrary model misspecification, and constructs large-sample confidence intervals for the average treatment effects. In addition, we also study the robust utilization of models for understanding treatment effect heterogeneity across individuals. In the last chapter, we focus on the unit allocation in fixed K-stages of a bandit experiment. Exploration-exploitation dilemma, the balance between exploring the environment to find the most profitable action arm and exploiting the best action arm based on the current understanding of the environment, is a problem in reinforcement learning. To study such a balance, the multi-armed bandit problem is a simple but essential model. The typical bandit is allocating units to arms and collecting the outcome one by one, which is powerful but can be time-consuming. In this chapter, we consider the setting where the experimentation for all units has to be completed in a fixed number of stages, where at each stage, multiple units will be allocated at the same time. We study the optimal way to allocate these units different stages. We propose a Bayesian approach as well as a Markov chain Monte Carlo method to find the ``optimal'' allocation."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Analysis based on incomplete data"]}]}],"canonical_facts":{"dc:contributor":["Li, Bo","Li, Xinran","Shao, Xiaofeng","Wang, Shulei"],"dc:creator":["Qu, Tianyi"],"dc:date":["2022-12","2022-11-29"],"dc:description":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2024-12-01","The student, Tianyi Qu, accepted the attached license on 2022-11-21 at 19:36.","The student, Tianyi Qu, submitted this Dissertation for approval on 2022-11-27 at 00:29.","This Dissertation was approved for publication on 2022-11-29 at 10:54.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18612 on 2023-04-12 at 11:35:07","In many situations, we may have to rely on partial information to do data analysis due to various reasons. For example, it may be that part of the data is not available, such as missing values and causal inference, or the analysis has to be drawn before the whole data is collected, such as bandit. In this thesis, we address three challenges that involve analysis with partial information. In the first chapter, we focus on prediction with unbalanced missing values. Public health data, such as HIV new diagnoses, are often left-censored due to confidentiality issues. Standard analysis approaches that assume censored values as missing at random often lead to biased estimates and inferior predictions. Motivated by the Philadelphia areal counts of HIV new diagnoses for which all values less than or equal to 5 are suppressed, we propose two methods to reduce the adverse influence of missingness on predictions and imputation of areal HIV new diagnoses. One is the likelihood-based method that integrates the missing mechanism into the likelihood function, and the other is a nonparametric algorithm for matrix factorization imputation. Numerical studies and the Philadelphia data analysis demonstrate that the two proposed methods can significantly improve prediction and imputation based on left-censored HIV data. We also compare the two methods on their robustness to model misspecification and find that both methods appear to be robust for prediction, while their performance for imputation depends on model specification. In the second chapter, we focus on the causal inference that for each individual, either treatment or control outcome can be observed. Randomized experiments have been the gold standard for drawing causal inferences. Conventional model-based analysis has been one of the most popular ways of analyzing treatment effects from randomized experiments, which is often carried through inference for certain model parameters. We provide a systematic investigation of model-based analysis, including the theory of M-estimation, under the randomization-based inference framework, avoiding any distributional assumptions on outcomes or covariates and utilizing only randomization as the ``reasoned basis''. We first show that the conventional model-based approach generally provides biased treatment effect estimation. We then study the model-imputed approach that uses the models mainly as a tool for imputing potential outcomes. Such an approach, although generally leading to biased estimation as well, can be valid for some special classes of models, e.g., the generalized linear model with canonical links. We finally recommend the model-assisted approach, which always provides consistent estimation and is robust to arbitrary model misspecification, and constructs large-sample confidence intervals for the average treatment effects. In addition, we also study the robust utilization of models for understanding treatment effect heterogeneity across individuals. In the last chapter, we focus on the unit allocation in fixed K-stages of a bandit experiment. Exploration-exploitation dilemma, the balance between exploring the environment to find the most profitable action arm and exploiting the best action arm based on the current understanding of the environment, is a problem in reinforcement learning. To study such a balance, the multi-armed bandit problem is a simple but essential model. The typical bandit is allocating units to arms and collecting the outcome one by one, which is powerful but can be time-consuming. In this chapter, we consider the setting where the experimentation for all units has to be completed in a fixed number of stages, where at each stage, multiple units will be allocated at the same time. We study the optimal way to allocate these units different stages. We propose a Bayesian approach as well as a Markov chain Monte Carlo method to find the ``optimal'' allocation."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/117562"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Tianyi Qu"],"dc:subject":["Likelihood","Matrix Factorization","Missing Value","Spatiotemporal Data","Randomization-based Inference","Potential Outcomes","M-estimation","Generalized Linear Models","Treatment Effect Heterogeneity","Model-assisted Approach","Bandit Problems","Adaptive Allocation","Metropolis-hasting Optimization"],"dc:title":["Analysis based on incomplete data"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:56Z"}