{"id":{"repo_id":"rockefeller","oai_identifier":"oai:digitalcommons.rockefeller.edu:student_theses_and_dissertations-1807"},"canonical_url":"https://search.dev.ndltd.org/etd/rockefeller/oai:digitalcommons.rockefeller.edu:student_theses_and_dissertations-1807","repository":{"repo_id":"rockefeller","name":"Rockefeller","base_url":"https://digitalcommons.rockefeller.edu/do/oai/"},"display":{"title":"Benchmarking Reinforcement Learning and Off Policy Evaluation for Medical Decision Making","abstract":"<p>Healthcare applications pose significant challenges to existing Reinforcement Learning (RL) methods due to implementation risks, low data availability, short treatment episodes, sparse re[1]wards, partial observations, and heterogeneous treatment effects (HTE). Despite significant interest in developing Dynamic Treatment Regimes (DTRs) for longitudinal patient care scenarios, no standardized benchmark has yet been developed. To address this gap, this thesis introduces Episodes of Care (EpiCare), a benchmark designed to mimic the challenges associated with applying RL to longitudinal healthcare settings. I leverage this benchmark to test seven state-of-the-art offline RL models as well as five common off-policy evaluation (OPE) techniques. My results suggest that while offline RL may be capable of improving upon existing standards of care given large data availability, its applicability does not appear to extend to the moderate to low data regimes typical of healthcare settings. Additionally, I demonstrate that several OPE techniques which have become standard in the medical RL literature fail to perform adequately under simulated conditions. These results suggest that the performance of RL models in DTRs may be difficult to meaningfully evaluate using current OPE methods, indicating that RL for this application may still be in its early stages. It is my hope that these findings, along with the EpiCare benchmark itself, will facilitate the comparison of existing methods and inspire further research into techniques that increase the practical applicability of medical RL.</p>","abstract_html":"&lt;p&gt;Healthcare applications pose significant challenges to existing Reinforcement Learning (RL) methods due to implementation risks, low data availability, short treatment episodes, sparse re[1]wards, partial observations, and heterogeneous treatment effects (HTE). Despite significant interest in developing Dynamic Treatment Regimes (DTRs) for longitudinal patient care scenarios, no standardized benchmark has yet been developed. To address this gap, this thesis introduces Episodes of Care (EpiCare), a benchmark designed to mimic the challenges associated with applying RL to longitudinal healthcare settings. I leverage this benchmark to test seven state-of-the-art offline RL models as well as five common off-policy evaluation (OPE) techniques. My results suggest that while offline RL may be capable of improving upon existing standards of care given large data availability, its applicability does not appear to extend to the moderate to low data regimes typical of healthcare settings. Additionally, I demonstrate that several OPE techniques which have become standard in the medical RL literature fail to perform adequately under simulated conditions. These results suggest that the performance of RL models in DTRs may be difficult to meaningfully evaluate using current OPE methods, indicating that RL for this application may still be in its early stages. It is my hope that these findings, along with the EpiCare benchmark itself, will facilitate the comparison of existing methods and inspire further research into techniques that increase the practical applicability of medical RL.&lt;/p&gt;","abstract_has_math":false,"creators":["Hargrave, Mason"],"institution":null,"degree_name":"Doctor of Philosophy (PhD)","degree_level":"Thesis","degree_discipline":null,"degree_department":null,"school":null,"contributors":["Marcelo O. Magnasco"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-01-01T08:00:00Z","date_published":"2025-01-01T08:00:00Z","updated_at":"2026-07-24T04:11:51Z","subjects":["reinforcement learning","dynamic treatment regimes","offline RL","healthcare","off-policy evaluation","benchmark dataset","Life Sciences"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://digitalcommons.rockefeller.edu/student_theses_and_dissertations/803","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Marcelo O. Magnasco"]},{"key":"dc:creator","label":"Author","values":["Hargrave, Mason"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["reinforcement learning","dynamic treatment regimes","offline RL","healthcare","off-policy evaluation","benchmark dataset","Life Sciences"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.rockefeller.edu/student_theses_and_dissertations/803"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Healthcare applications pose significant challenges to existing Reinforcement Learning (RL) methods due to implementation risks, low data availability, short treatment episodes, sparse re[1]wards, partial observations, and heterogeneous treatment effects (HTE). Despite significant interest in developing Dynamic Treatment Regimes (DTRs) for longitudinal patient care scenarios, no standardized benchmark has yet been developed. To address this gap, this thesis introduces Episodes of Care (EpiCare), a benchmark designed to mimic the challenges associated with applying RL to longitudinal healthcare settings. I leverage this benchmark to test seven state-of-the-art offline RL models as well as five common off-policy evaluation (OPE) techniques. My results suggest that while offline RL may be capable of improving upon existing standards of care given large data availability, its applicability does not appear to extend to the moderate to low data regimes typical of healthcare settings. Additionally, I demonstrate that several OPE techniques which have become standard in the medical RL literature fail to perform adequately under simulated conditions. These results suggest that the performance of RL models in DTRs may be difficult to meaningfully evaluate using current OPE methods, indicating that RL for this application may still be in its early stages. It is my hope that these findings, along with the EpiCare benchmark itself, will facilitate the comparison of existing methods and inspire further research into techniques that increase the practical applicability of medical RL.</p>"]},{"key":"dc:title","label":"Title","values":["Benchmarking Reinforcement Learning and Off Policy Evaluation for Medical Decision Making"]}]}],"canonical_facts":{"dc:contributor":["Marcelo O. Magnasco"],"dc:creator":["Hargrave, Mason"],"dc:description.abstract":["<p>Healthcare applications pose significant challenges to existing Reinforcement Learning (RL) methods due to implementation risks, low data availability, short treatment episodes, sparse re[1]wards, partial observations, and heterogeneous treatment effects (HTE). Despite significant interest in developing Dynamic Treatment Regimes (DTRs) for longitudinal patient care scenarios, no standardized benchmark has yet been developed. To address this gap, this thesis introduces Episodes of Care (EpiCare), a benchmark designed to mimic the challenges associated with applying RL to longitudinal healthcare settings. I leverage this benchmark to test seven state-of-the-art offline RL models as well as five common off-policy evaluation (OPE) techniques. My results suggest that while offline RL may be capable of improving upon existing standards of care given large data availability, its applicability does not appear to extend to the moderate to low data regimes typical of healthcare settings. Additionally, I demonstrate that several OPE techniques which have become standard in the medical RL literature fail to perform adequately under simulated conditions. These results suggest that the performance of RL models in DTRs may be difficult to meaningfully evaluate using current OPE methods, indicating that RL for this application may still be in its early stages. It is my hope that these findings, along with the EpiCare benchmark itself, will facilitate the comparison of existing methods and inspire further research into techniques that increase the practical applicability of medical RL.</p>"],"dc:identifier":["https://digitalcommons.rockefeller.edu/student_theses_and_dissertations/803"],"dc:subject":["reinforcement learning","dynamic treatment regimes","offline RL","healthcare","off-policy evaluation","benchmark dataset","Life Sciences"],"dc:title":["Benchmarking Reinforcement Learning and Off Policy Evaluation for Medical Decision Making"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-24T04:11:51Z"}