Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 8 of 8 for “"offline RL"”.

  1. Quasimetric decision transformer: enhancing goal-conditioned reinforcement learning with structured distance guidance

    Recent works have shown that tackling offline Reinforcement Learning (RL) with a conditional policy produces promising results. Decision Transformer (DT) have shown promising results in offline RL by leveraging sequence modeling. However, standard DTs rely on Returns-to-Go (RTG) tokens, which are …

    uoit Repository record for Quasimetric decision transformer: enhancing goal-conditioned reinforcement learning with structured distance guidance (opens in a new tab)

  2. Towards Fine-grained Multi-Attribute Control using Language Models

    … toxicity detection. Next, I introduce a novel offline RL algorithm that can utilize arbitrary numeric scores as rewards during training to optimize any user-desired LM behavior by filtering out suboptimal data. Finally, I designed an offline RL framework, I propose a fine-grained …

    gatech Repository record for Towards Fine-grained Multi-Attribute Control using Language Models (opens in a new tab)

  3. Benchmarking Reinforcement Learning and Off Policy Evaluation for Medical Decision Making

    … challenges to existing Reinforcement Learning (RL) methods due to implementation risks, low data availability, short treatment episodes, sparse re[1]wards, partial observations, and heterogeneous treatment effects (HTE). Despite significant interest in developing Dynamic Treatment Regimes (DTRs) …

    rockefeller Repository record for Benchmarking Reinforcement Learning and Off Policy Evaluation for Medical Decision Making (opens in a new tab)

  4. Data-Efficient Offline Reinforcement Learning with Heterogeneous Agents

    Performance of state-of-the art offline and model-based reinforcement learning (RL) algorithms deteriorates significantly when subjected to severe data scarcity and the presence of heterogeneous agents. In this work, we propose a model-based offline RL method to approach this setting. Using all …

    mit Repository record for Data-Efficient Offline Reinforcement Learning with Heterogeneous Agents (opens in a new tab)

  5. Generative Discovery via Reinforcement Learning

    … design, etc.). Reinforcement learning (RL) is well-suited for discovery tasks because it enables machines to learn through trial and error. My work overcomes the following major limitation of today’s RL algorithms and thereby advances their discovery potential: Mitigate the bias of …

    mit Repository record for Generative Discovery via Reinforcement Learning (opens in a new tab)

  6. PerSim: Data-Efficient Offline Reinforcement Learning with Heterogeneous Agents via Latent Factor Representation

    Offline reinforcement learning, where a policy is learned from a fixed dataset of trajectories without further interaction with the environment, is one of the greatest challenges in reinforcement learning. Despite its compelling application to large, real-world datasets, existing RL benchmarks have …

    mit Repository record for PerSim: Data-Efficient Offline Reinforcement Learning with Heterogeneous Agents via Latent Factor Representation (opens in a new tab)

  7. Learning value functions from undirected state-only experience

    Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-04-12 without embargo terms

    uiuc Repository record for Learning value functions from undirected state-only experience (opens in a new tab)

  8. Batch value function tournament for offline policy selection in reinforcement learning

    Offline policy selection is a challenging open problem in reinforcement learning that has many important applications. The recently proposed Batch Value Function Tournament (BVFT) algorithm for batch learning offers some nice properties and can be applied to the model selection problem. In this …

    uiuc Repository record for Batch value function tournament for offline policy selection in reinforcement learning (opens in a new tab)