Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 38 for “"Reward functions"”.
-
Rank2Reward: Learning Robot Reward Functions from Passive Video
… states and actions is by inferring a well-shaped reward function for reinforcement learning. The challenging problem is determining how to ground visual demonstration inputs into a well-shaped and informative reward function for reinforcement learning. To this end, we propose a technique, …
-
Loss and Reward Functions for Generative Question Answering Systems
Recent advancements in AI, mainly through Large Language Models (LLMs), have transformed the field, driving both industrial and academic progress. These models are typically trained using causal language modeling tasks, which, while effective, face certain limitations compared to traditional …
-
Learning Heterogeneous Resource-Constrained Task Allocation Using Concurrent Multi-Task Bandits
… often assume that task requirements or reward functions are known and explicitly specified by the user in advance. In this thesis, we explore the challenge of forming effective coalitions for a given heterogeneous multi-robot team when task reward functions are unknown. To tackle this …
-
Tame Long-Horizon Model-Based Reinforcement Learning
… that naturally enables transfer across different reward functions, but struggles to scale to complex environments due to the compounding error of applying a learned dynamics model iteratively. To get the best of both worlds, we propose a self-supervised reinforcement learning method that enables …
-
Bayesian nonparametric reward learning from demonstration
… inverse reinforcement learning, whereby the reward of the demonstrator is inferred. Formally, inverse reinforcement learning (IRL) is the task of learning the reward function of a Markov Decision Process (MDP) given knowledge of the transition function and a set of observed demonstrations. …
-
Computationally efficient Gaussian Process changepoint detection and regression
… for multi-task Reinforcement Learning when the reward function is nonstationary. First, a novel algorithm UCRL-GP is introduced for stationary reward functions. Then, UCRL-GP is combined with GP-NBC to create UCRL-GP-CPD, which is an algorithm for nonstationary reward functions. Unlike previous …
-
On the Learnability of General Reinforcement-Learning Objectives
… these objectives are expressed through reward functions, enabling well-established guarantees on learning near-optimal policies with a high probability — a property known as probably approximately correct (PAC) -learnability. However, reward functions often serve as imperfect surrogates …
-
Improving service level agreements for a job scheduler by visualizing simulations
… that both helps users visualize SLAs and their reward functions, and allows users to create an SLA and gain an idea of the behavior of a job scheduler with the SLA as input.
-
Goal alignment: re-analyzing value alignment problems using human-aware AI
… specification mechanism, for example, the use of reward functions. However, the complexity of the objective specification mechanism is just one of many reasons why the user may have misspecified their objective. A foundational cause for misspecification that is being overlooked by the previous …
-
Model-free reinforcement learning in non-stationary Markov Decision Processes
… problem where an agent maximizes its cumulative reward through sequential interactions with an initially unknown environment, usually modeled by a Markov Decision Process (MDP). The classical RL literature typically assumes that the state transition functions and the reward functions of the MDP …
-
Reinforcement Learning for Self-adapting Time Discretizations of Complex Systems
… beds, this research develops models to determine reward functions and dynamically tunes controller parameters that minimize both the error and number of steps required for approximate mathematical solutions. Our best reward function is based on an error that does not overly punish rejected states. …
-
Reactive, Autonomous, Markovian Sensor Tasking in Communication Starved Environments
… used as the simulation framework to test various reward functions and decision algorithms while enabling autonomous, reactive sensor tasking. The goal of this work was used the developed evaluation methodology to perform statistical analyses to determine which metrics were most reliable and …
-
Autonomous Source Localization
… characterized the performance of several reward functions and different exploration algorithms in scenarios covering a range of source strengths and region sizes. These experiments demonstrated the improved performance of planning-based algorithms over the myopic method initially tested in …
-
Irreversible Actions in Assistance Games with a Dynamic Goal
Reinforcement Learning (RL) agents optimize reward functions to learn desirable policies in a variety of important real-world applications such as self-driving cars and recommender systems. However, in practice, it can be very difficult to specify the correct reward function for a complex problem, …
-
Future of Personalized, Aligned Language Models
… and the policy. We present a new framework for reward optimization, Value Augmented Sampling (VAS), that can maximize different reward functions using data sampled from only the initial, frozen LLM. VAS solves for the optimal reward-maximizing policy without co-training the policy and the value …
-
Baby Gym: Bridging the Gap between Reinforcement Learning and Human Infant Locomotor Development
… a reproducible RL environment with several new reward functions that yield human-like locomotor development stages; and initial methods for evaluating the "human-likeness" of the emerged locomotion.
-
Prompt Injection Generation Using Small Language Models with Reinforcement Learning with Artificial Intelligence Feedback
… instead. By optimizing the reference model and reward functions, we improve alignment with ground truth prompt injection messages while addressing issues such as mode collapse and overfitting. These findings show promise, and further research is necessary to determine how well the approach can …
-
Adaptive Collaborative Channel Finding Approaches for Autonomous Marine Vehicles
… process (MDP) planning with two state-of-the-art reward functions: Upper Confidence Bound (UCB) and Maximum Value Information (MVI). The performance of each method is evaluated through comparison of the time it takes to identify a continuous channel through an area, using one, two, three, or four …
-
Virtual humans making first contact : teaching socially appropriate approaching behavior using deep reinforcement learning
… in Unity and experimenting with different reward functions, virtual humans (agents) are trained to approach another virtual human (target), while adhering to this model as much as possible. The results are assessed numerically, as well as visually. Finally, a perceptual study compares the …
-
Optimal Reinforcement Learning with Black Holes
… in which we lose all turn information and all reward from trajectories that visit a particular subset of states. We assume awareness of the trajectory loss events, making this a censored data problem but not a truncated data problem. We have elected to work in a fixed-horizon setting for easier …
Page 1 of 2