Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 8 of 8 for “"offline RL"”.
-
Quasimetric decision transformer: enhancing goal-conditioned reinforcement learning with structured distance guidance
Recent works have shown that tackling offline Reinforcement Learning (RL) with a conditional policy produces promising results. Decision Transformer (DT) have shown promising results in offline RL by leveraging sequence modeling. However, standard DTs rely on Returns-to-Go (RTG) tokens, which are …
-
Towards Fine-grained Multi-Attribute Control using Language Models
… toxicity detection. Next, I introduce a novel offline RL algorithm that can utilize arbitrary numeric scores as rewards during training to optimize any user-desired LM behavior by filtering out suboptimal data. Finally, I designed an offline RL framework, I propose a fine-grained …
-
Benchmarking Reinforcement Learning and Off Policy Evaluation for Medical Decision Making
… challenges to existing Reinforcement Learning (RL) methods due to implementation risks, low data availability, short treatment episodes, sparse re[1]wards, partial observations, and heterogeneous treatment effects (HTE). Despite significant interest in developing Dynamic Treatment Regimes (DTRs) …
-
Data-Efficient Offline Reinforcement Learning with Heterogeneous Agents
Performance of state-of-the art offline and model-based reinforcement learning (RL) algorithms deteriorates significantly when subjected to severe data scarcity and the presence of heterogeneous agents. In this work, we propose a model-based offline RL method to approach this setting. Using all …
-
Generative Discovery via Reinforcement Learning
… design, etc.). Reinforcement learning (RL) is well-suited for discovery tasks because it enables machines to learn through trial and error. My work overcomes the following major limitation of today’s RL algorithms and thereby advances their discovery potential: Mitigate the bias of …
-
PerSim: Data-Efficient Offline Reinforcement Learning with Heterogeneous Agents via Latent Factor Representation
Offline reinforcement learning, where a policy is learned from a fixed dataset of trajectories without further interaction with the environment, is one of the greatest challenges in reinforcement learning. Despite its compelling application to large, real-world datasets, existing RL benchmarks have …
-
Learning value functions from undirected state-only experience
Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-04-12 without embargo terms
-
Batch value function tournament for offline policy selection in reinforcement learning
Offline policy selection is a challenging open problem in reinforcement learning that has many important applications. The recently proposed Batch Value Function Tournament (BVFT) algorithm for batch learning offers some nice properties and can be applied to the model selection problem. In this …