{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/116243"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/116243","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Statistical reinforcement learning for individualized decision making","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2022-11-15 without embargo terms","abstract_has_math":false,"creators":["Zhou, Wenzhuo"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Statistics","degree_department":null,"school":null,"contributors":["Zhu, Ruoqing","Qu, Annie","Shao, Xiaofeng","Li, Xinran"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-08","date_published":"2022-08","updated_at":"2026-07-22T22:24:55Z","subjects":["Dynamic Treatment Regimes","Markov Decision Process","Reinforcement Learning","Survival Analysis","Dimension Reduction"],"languages":["en","eng"],"rights":["Copyright 2022 Wenzhuo Zhou"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/116243","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhu, Ruoqing","Qu, Annie","Shao, Xiaofeng","Li, Xinran"]},{"key":"dc:creator","label":"Author","values":["Zhou, Wenzhuo"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-08","2022-07-15"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Dynamic Treatment Regimes","Markov Decision Process","Reinforcement Learning","Survival Analysis","Dimension Reduction"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Wenzhuo Zhou"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/116243"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","The student, Wenzhuo Zhou, accepted the attached license on 2022-07-14 at 11:35.","The student, Wenzhuo Zhou, submitted this Dissertation for approval on 2022-07-14 at 11:59.","This Dissertation was approved for publication on 2022-07-15 at 08:58.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18307 on 2022-11-15 at 18:21:04","Learning personalized patterns from highly heterogeneous data for decision making is one of the most challenging tasks in modern machine learning. However, it is important and necessary for the successes of many real-world applications, such as personalized medicine, mobile health, robotics, ride-sharing, etc. In general, this decision-making procedure can be formally characterized by reinforcement learning, or say dynamic treatment regime if it is associated with healthcare domain. In this thesis, we address several major challenges for individualized decision making in reinforcement learning and dynamic treatment regimes. On the perspective of the number of decision stages, we tackle the problems over finite and infinite length of horizon settings. On the action space point of view, we propose the methods which can be applicable for discrete cases, e.g., binary, multi-category and also for continuous cases. In the first part of the thesis, we focus on developing a novel personalized decision making procedure for dose assignment problems. Learning an individualized dose rule in personalized medicine is a challenging statistical problem. Existing methods often suffer from the curse of dimensionality, especially when the decision function is estimated nonparametrically. To tackle this problem, we propose a dimension reduction framework that effectively reduces the estimation to a lower-dimensional subspace of the covariates. We exploit that the individualized dose rule can be defined in a subspace spanned by a few linear combinations of the covariates, leading to a more parsimonious model. The proposed framework does not require the inverse probability of the propensity score under observational studies due to a direct maximization of the value function. Under the same framework, we further propose a pseudo-direct learning approach that focuses more on estimating the dimensionality-reduced subspace of the treatment outcome. Parameters in both approaches can be estimated efficiently using an orthogonality constrained optimization algorithm on the Stiefel manifold. Under mild regularity assumptions, the results on the asymptotic normality and consistency of the proposed estimators are established, respectively. In the second part, we develop a novel angle-based approach to search the optimal dynamic treatment regime (DTR) under a multicategory treatment framework for survival data. The proposed method targets to maximize the conditional survival function of patients following a DTR. Specifically, the proposed method obtains the optimal DTR via integrating estimations of decision rules at multiple stages into a single multicategory classification algorithm without imposing additional constraints, which is also more computationally efficient and robust. In theory, we establish Fisher consistency and provide the risk bound for the proposed estimator under regularity conditions. In the third component of the thesis, we tackle some fundamental challenges for reinforcement learning and mobile health. The practical use of mobile health (mHealth) technology raises unique challenges to existing methodologies on learning an optimal dynamic treatment regime. Many mHealth applications involve decision-making with large numbers of intervention options and under an infinite time horizon setting where the number of decision stages diverges to infinity. In addition, temporary medication shortages may cause optimal treatments to be unavailable, while it is unclear what alternatives can be used. To address these challenges, we propose a proximal temporal consistency learning framework to estimate an optimal regime that is adaptively adjusted between deterministic and stochastic sparse policy models. The resulting minimax estimator avoids the double sampling issue in the existing algorithms. It can be further simplified and can easily incorporate off-policy data without mismatched distribution corrections. We study theoretical properties of the sparse policy and establish finite-sample bounds on the excess risk and performance error."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Statistical reinforcement learning for individualized decision making"]}]}],"canonical_facts":{"dc:contributor":["Zhu, Ruoqing","Qu, Annie","Shao, Xiaofeng","Li, Xinran"],"dc:creator":["Zhou, Wenzhuo"],"dc:date":["2022-08","2022-07-15"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","The student, Wenzhuo Zhou, accepted the attached license on 2022-07-14 at 11:35.","The student, Wenzhuo Zhou, submitted this Dissertation for approval on 2022-07-14 at 11:59.","This Dissertation was approved for publication on 2022-07-15 at 08:58.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18307 on 2022-11-15 at 18:21:04","Learning personalized patterns from highly heterogeneous data for decision making is one of the most challenging tasks in modern machine learning. However, it is important and necessary for the successes of many real-world applications, such as personalized medicine, mobile health, robotics, ride-sharing, etc. In general, this decision-making procedure can be formally characterized by reinforcement learning, or say dynamic treatment regime if it is associated with healthcare domain. In this thesis, we address several major challenges for individualized decision making in reinforcement learning and dynamic treatment regimes. On the perspective of the number of decision stages, we tackle the problems over finite and infinite length of horizon settings. On the action space point of view, we propose the methods which can be applicable for discrete cases, e.g., binary, multi-category and also for continuous cases. In the first part of the thesis, we focus on developing a novel personalized decision making procedure for dose assignment problems. Learning an individualized dose rule in personalized medicine is a challenging statistical problem. Existing methods often suffer from the curse of dimensionality, especially when the decision function is estimated nonparametrically. To tackle this problem, we propose a dimension reduction framework that effectively reduces the estimation to a lower-dimensional subspace of the covariates. We exploit that the individualized dose rule can be defined in a subspace spanned by a few linear combinations of the covariates, leading to a more parsimonious model. The proposed framework does not require the inverse probability of the propensity score under observational studies due to a direct maximization of the value function. Under the same framework, we further propose a pseudo-direct learning approach that focuses more on estimating the dimensionality-reduced subspace of the treatment outcome. Parameters in both approaches can be estimated efficiently using an orthogonality constrained optimization algorithm on the Stiefel manifold. Under mild regularity assumptions, the results on the asymptotic normality and consistency of the proposed estimators are established, respectively. In the second part, we develop a novel angle-based approach to search the optimal dynamic treatment regime (DTR) under a multicategory treatment framework for survival data. The proposed method targets to maximize the conditional survival function of patients following a DTR. Specifically, the proposed method obtains the optimal DTR via integrating estimations of decision rules at multiple stages into a single multicategory classification algorithm without imposing additional constraints, which is also more computationally efficient and robust. In theory, we establish Fisher consistency and provide the risk bound for the proposed estimator under regularity conditions. In the third component of the thesis, we tackle some fundamental challenges for reinforcement learning and mobile health. The practical use of mobile health (mHealth) technology raises unique challenges to existing methodologies on learning an optimal dynamic treatment regime. Many mHealth applications involve decision-making with large numbers of intervention options and under an infinite time horizon setting where the number of decision stages diverges to infinity. In addition, temporary medication shortages may cause optimal treatments to be unavailable, while it is unclear what alternatives can be used. To address these challenges, we propose a proximal temporal consistency learning framework to estimate an optimal regime that is adaptively adjusted between deterministic and stochastic sparse policy models. The resulting minimax estimator avoids the double sampling issue in the existing algorithms. It can be further simplified and can easily incorporate off-policy data without mismatched distribution corrections. We study theoretical properties of the sparse policy and establish finite-sample bounds on the excess risk and performance error."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/116243"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Wenzhuo Zhou"],"dc:subject":["Dynamic Treatment Regimes","Markov Decision Process","Reinforcement Learning","Survival Analysis","Dimension Reduction"],"dc:title":["Statistical reinforcement learning for individualized decision making"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:55Z"}