Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 38 for “"Reward functions"”.

  1. Rank2Reward: Learning Robot Reward Functions from Passive Video

    … states and actions is by inferring a well-shaped reward function for reinforcement learning. The challenging problem is determining how to ground visual demonstration inputs into a well-shaped and informative reward function for reinforcement learning. To this end, we propose a technique, …

    mit Repository record for Rank2Reward: Learning Robot Reward Functions from Passive Video (opens in a new tab)

  2. Loss and Reward Functions for Generative Question Answering Systems

    Recent advancements in AI, mainly through Large Language Models (LLMs), have transformed the field, driving both industrial and academic progress. These models are typically trained using causal language modeling tasks, which, while effective, face certain limitations compared to traditional …

    trento Repository record for Loss and Reward Functions for Generative Question Answering Systems (opens in a new tab)

  3. Learning Heterogeneous Resource-Constrained Task Allocation Using Concurrent Multi-Task Bandits

    … often assume that task requirements or reward functions are known and explicitly specified by the user in advance. In this thesis, we explore the challenge of forming effective coalitions for a given heterogeneous multi-robot team when task reward functions are unknown. To tackle this …

    gatech Repository record for Learning Heterogeneous Resource-Constrained Task Allocation Using Concurrent Multi-Task Bandits (opens in a new tab)

  4. Tame Long-Horizon Model-Based Reinforcement Learning

    … that naturally enables transfer across different reward functions, but struggles to scale to complex environments due to the compounding error of applying a learned dynamics model iteratively. To get the best of both worlds, we propose a self-supervised reinforcement learning method that enables …

    mit Repository record for Tame Long-Horizon Model-Based Reinforcement Learning (opens in a new tab)

  5. Bayesian nonparametric reward learning from demonstration

    … inverse reinforcement learning, whereby the reward of the demonstrator is inferred. Formally, inverse reinforcement learning (IRL) is the task of learning the reward function of a Markov Decision Process (MDP) given knowledge of the transition function and a set of observed demonstrations. …

    mit Repository record for Bayesian nonparametric reward learning from demonstration (opens in a new tab)

  6. Computationally efficient Gaussian Process changepoint detection and regression

    … for multi-task Reinforcement Learning when the reward function is nonstationary. First, a novel algorithm UCRL-GP is introduced for stationary reward functions. Then, UCRL-GP is combined with GP-NBC to create UCRL-GP-CPD, which is an algorithm for nonstationary reward functions. Unlike previous …

    mit Repository record for Computationally efficient Gaussian Process changepoint detection and regression (opens in a new tab)

  7. On the Learnability of General Reinforcement-Learning Objectives

    … these objectives are expressed through reward functions, enabling well-established guarantees on learning near-optimal policies with a high probability — a property known as probably approximately correct (PAC) -learnability. However, reward functions often serve as imperfect surrogates …

    mit Repository record for On the Learnability of General Reinforcement-Learning Objectives (opens in a new tab)

  8. Improving service level agreements for a job scheduler by visualizing simulations

    … that both helps users visualize SLAs and their reward functions, and allows users to create an SLA and gain an idea of the behavior of a job scheduler with the SLA as input.

    mit Repository record for Improving service level agreements for a job scheduler by visualizing simulations (opens in a new tab)

  9. Goal alignment: re-analyzing value alignment problems using human-aware AI

    … specification mechanism, for example, the use of reward functions. However, the complexity of the objective specification mechanism is just one of many reasons why the user may have misspecified their objective. A foundational cause for misspecification that is being overlooked by the previous …

    colostate Repository record for Goal alignment: re-analyzing value alignment problems using human-aware AI (opens in a new tab)

  10. Model-free reinforcement learning in non-stationary Markov Decision Processes

    … problem where an agent maximizes its cumulative reward through sequential interactions with an initially unknown environment, usually modeled by a Markov Decision Process (MDP). The classical RL literature typically assumes that the state transition functions and the reward functions of the MDP …

    uiuc Repository record for Model-free reinforcement learning in non-stationary Markov Decision Processes (opens in a new tab)

  11. Reinforcement Learning for Self-adapting Time Discretizations of Complex Systems

    … beds, this research develops models to determine reward functions and dynamically tunes controller parameters that minimize both the error and number of steps required for approximate mathematical solutions. Our best reward function is based on an error that does not overly punish rejected states. …

    vt Repository record for Reinforcement Learning for Self-adapting Time Discretizations of Complex Systems (opens in a new tab)

  12. Reactive, Autonomous, Markovian Sensor Tasking in Communication Starved Environments

    … used as the simulation framework to test various reward functions and decision algorithms while enabling autonomous, reactive sensor tasking. The goal of this work was used the developed evaluation methodology to perform statistical analyses to determine which metrics were most reliable and …

    vt Repository record for Reactive, Autonomous, Markovian Sensor Tasking in Communication Starved Environments (opens in a new tab)

  13. Autonomous Source Localization

    … characterized the performance of several reward functions and different exploration algorithms in scenarios covering a range of source strengths and region sizes. These experiments demonstrated the improved performance of planning-based algorithms over the myopic method initially tested in …

    vt Repository record for Autonomous Source Localization (opens in a new tab)

  14. Irreversible Actions in Assistance Games with a Dynamic Goal

    Reinforcement Learning (RL) agents optimize reward functions to learn desirable policies in a variety of important real-world applications such as self-driving cars and recommender systems. However, in practice, it can be very difficult to specify the correct reward function for a complex problem, …

    mit Repository record for Irreversible Actions in Assistance Games with a Dynamic Goal (opens in a new tab)

  15. Future of Personalized, Aligned Language Models

    … and the policy. We present a new framework for reward optimization, Value Augmented Sampling (VAS), that can maximize different reward functions using data sampled from only the initial, frozen LLM. VAS solves for the optimal reward-maximizing policy without co-training the policy and the value …

    mit Repository record for Future of Personalized, Aligned Language Models (opens in a new tab)

  16. Baby Gym: Bridging the Gap between Reinforcement Learning and Human Infant Locomotor Development

    … a reproducible RL environment with several new reward functions that yield human-like locomotor development stages; and initial methods for evaluating the "human-likeness" of the emerged locomotion.

    mit Repository record for Baby Gym: Bridging the Gap between Reinforcement Learning and Human Infant Locomotor Development (opens in a new tab)

  17. Prompt Injection Generation Using Small Language Models with Reinforcement Learning with Artificial Intelligence Feedback

    … instead. By optimizing the reference model and reward functions, we improve alignment with ground truth prompt injection messages while addressing issues such as mode collapse and overfitting. These findings show promise, and further research is necessary to determine how well the approach can …

    mit Repository record for Prompt Injection Generation Using Small Language Models with Reinforcement Learning with Artificial Intelligence Feedback (opens in a new tab)

  18. Adaptive Collaborative Channel Finding Approaches for Autonomous Marine Vehicles

    … process (MDP) planning with two state-of-the-art reward functions: Upper Confidence Bound (UCB) and Maximum Value Information (MVI). The performance of each method is evaluated through comparison of the time it takes to identify a continuous channel through an area, using one, two, three, or four …

    mit Repository record for Adaptive Collaborative Channel Finding Approaches for Autonomous Marine Vehicles (opens in a new tab)

  19. Virtual humans making first contact : teaching socially appropriate approaching behavior using deep reinforcement learning

    … in Unity and experimenting with different reward functions, virtual humans (agents) are trained to approach another virtual human (target), while adhering to this model as much as possible. The results are assessed numerically, as well as visually. Finally, a perceptual study compares the …

    reykjavik Repository record for Virtual humans making first contact : teaching socially appropriate approaching behavior using deep reinforcement learning (opens in a new tab)

  20. Optimal Reinforcement Learning with Black Holes

    … in which we lose all turn information and all reward from trajectories that visit a particular subset of states. We assume awareness of the trajectory loss events, making this a censored data problem but not a truncated data problem. We have elected to work in a fixed-horizon setting for easier …

    mit Repository record for Optimal Reinforcement Learning with Black Holes (opens in a new tab)

Page 1 of 2