{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/130009"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/130009","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Topics in offline statistical reinforcement learning: addressing challenges in continuous actions, distribution shifts, and unmeasured confounding","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-08-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2027-08-01","abstract_has_math":false,"creators":["Li, Yuhan"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Statistics","degree_department":null,"school":null,"contributors":["Zhu, Ruoqing","Shao, Xiaofeng","Zhao, Sihai Dave","Park, Chan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-06-19","date_published":"2025-06-19","updated_at":"2026-07-22T22:25:06Z","subjects":["Reinforcement Learning","Personalized Medicine","Causal Inference","Markov Decision Process","Policy Learning","Policy Evaluation"],"languages":["en","eng"],"rights":["Copyright 2025 Yuhan Li"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/130009","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhu, Ruoqing","Shao, Xiaofeng","Zhao, Sihai Dave","Park, Chan"]},{"key":"dc:creator","label":"Author","values":["Li, Yuhan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-06-19","2025-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Reinforcement Learning","Personalized Medicine","Causal Inference","Markov Decision Process","Policy Learning","Policy Evaluation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Yuhan Li"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/130009"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-08-01","The student, Yuhan Li, accepted the attached license on 2025-06-16 at 14:00.","The student, Yuhan Li, submitted this Dissertation for approval on 2025-06-16 at 14:16.","This Dissertation was approved for publication on 2025-06-19 at 13:08.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22339 on 2025-10-21 at 10:05:26","Reinforcement learning (RL) provides a principled framework for tackling sequential decision-making problems when system dynamics and outcomes are uncertain. While RL has made significant progress in recent decades, deploying these methods in real-world scenarios remain challenging. A major obstacle is that standard RL algorithms typically focuses on the online setting, where an agent continuously interacts with the environment to collect data, updating its policy in real time and learning by trial and error. In many practical domains, however, data collection is costly, and unconstrained exploration can raise serious safety and ethical concerns, especially in safety-critical areas such as personalized medicine and autonomous driving. Consequently, there is growing interest in offline RL, where the goal is to evaluate and optimize policies using only a fixed, precollected dataset, without any further interaction with the environment. In this thesis, we aim at tackling several major challenges in offline reinforcement learning. In the first part of the thesis, we focus on policy learning with continuous action space and introduce a novel quasi-optimal Bellman operator, which is able to identify near-optimal action regions. The proposed quasi-optimal Bellman operator addressed the shortcomings of existing approaches relying on modeling an optimal policy with infinite support distributions and is highly desirable in safety-critical scenarios. For the second part of this thesis, we study high-confidence off-policy evaluation in the context of infinite-horizon Markov decision processes, where the objective is to establish a confidence interval (CI) for the target policy value using only pre-collected data generated from unknown behavior policies. The proposed unified error-quantification framework handles distributional shift and balances the trade-off between bias and uncertainty to produce tight confidence intervals. The thrid part of the thesis considers offline policy learning with the existence of unmeausured confounder. We extend the proximal causal inference framework to infinite horizon and develop a novel identification results that enable nonparamatric estimation of policy value. Leveraging this identification result, we further develop a policy-gradient-type algorithm for offline policy learning despite hidden confounders."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Topics in offline statistical reinforcement learning: addressing challenges in continuous actions, distribution shifts, and unmeasured confounding"]}]}],"canonical_facts":{"dc:contributor":["Zhu, Ruoqing","Shao, Xiaofeng","Zhao, Sihai Dave","Park, Chan"],"dc:creator":["Li, Yuhan"],"dc:date":["2025-06-19","2025-08"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-08-01","The student, Yuhan Li, accepted the attached license on 2025-06-16 at 14:00.","The student, Yuhan Li, submitted this Dissertation for approval on 2025-06-16 at 14:16.","This Dissertation was approved for publication on 2025-06-19 at 13:08.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22339 on 2025-10-21 at 10:05:26","Reinforcement learning (RL) provides a principled framework for tackling sequential decision-making problems when system dynamics and outcomes are uncertain. While RL has made significant progress in recent decades, deploying these methods in real-world scenarios remain challenging. A major obstacle is that standard RL algorithms typically focuses on the online setting, where an agent continuously interacts with the environment to collect data, updating its policy in real time and learning by trial and error. In many practical domains, however, data collection is costly, and unconstrained exploration can raise serious safety and ethical concerns, especially in safety-critical areas such as personalized medicine and autonomous driving. Consequently, there is growing interest in offline RL, where the goal is to evaluate and optimize policies using only a fixed, precollected dataset, without any further interaction with the environment. In this thesis, we aim at tackling several major challenges in offline reinforcement learning. In the first part of the thesis, we focus on policy learning with continuous action space and introduce a novel quasi-optimal Bellman operator, which is able to identify near-optimal action regions. The proposed quasi-optimal Bellman operator addressed the shortcomings of existing approaches relying on modeling an optimal policy with infinite support distributions and is highly desirable in safety-critical scenarios. For the second part of this thesis, we study high-confidence off-policy evaluation in the context of infinite-horizon Markov decision processes, where the objective is to establish a confidence interval (CI) for the target policy value using only pre-collected data generated from unknown behavior policies. The proposed unified error-quantification framework handles distributional shift and balances the trade-off between bias and uncertainty to produce tight confidence intervals. The thrid part of the thesis considers offline policy learning with the existence of unmeausured confounder. We extend the proximal causal inference framework to infinite horizon and develop a novel identification results that enable nonparamatric estimation of policy value. Leveraging this identification result, we further develop a policy-gradient-type algorithm for offline policy learning despite hidden confounders."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/130009"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Yuhan Li"],"dc:subject":["Reinforcement Learning","Personalized Medicine","Causal Inference","Markov Decision Process","Policy Learning","Policy Evaluation"],"dc:title":["Topics in offline statistical reinforcement learning: addressing challenges in continuous actions, distribution shifts, and unmeasured confounding"],"dc:type":["text"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:06Z"}