Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 15 of 15 for “"Reward Shaping"”.
-
Theory and Application of Reward Shaping in Reinforcement Learning
… predictions; most notably that reducing the reward horizon implies faster learning. The bipedal walking task demonstrates that our reward shaping techniques allow a conventional reinforcement learning algorithm to find a good behavior efficiently despite a large state space with stochastic …
-
Efficient reinforcement learning via singular value decomposition, end-to-end model-based methods and reward shaping
… heuristics through the use of potential-based reward shaping and lifted function approximation.
-
Knowledge and Ignorance in Reinforcement Learning
… three: the state-space/transition model, the reward function, and the observation model. In this thesis, I present a framework for studying how the state of knowledge or uncertainty of each component affects the Reinforcement Learning process. I focus on the reward function and the observation …
-
Improving Exploration in Reinforcement Learning through Domain Knowledge and Parameter Analysis
… actions in order to maximise the long term reward. Solving this type of problems is a hard task when real domains of realistic size are considered because the state space grows exponentially with each state feature added to the representation of the problem. In its basic form, reinforcement …
-
Reinforcement Learning for Autonomous Aircraft Control and Aerial Maneuvering in Simulated Environments
… control. A combination of environment design and reward shaping was essential for achieving stable learning. Results show that performance improves significantly with training and proper reward design. Experiments across different difficulty levels highlight the impact of environment complexity on …
-
Rational communication for the coordination of multi-Agent systems
… from the field of Machine Learning, namely Reward Shaping, which allows the decentralised POMDP to be transformed into individual agent POMDPs that can be solved more easily. This approach can use a heuristic transformation to allow the approach to work in large problems like RobocupRescue …
-
Reinforcement Learning Control for Mobile Robot Parking with Safety Constraints
… safe set in real time. Experiments show that (i) reward shaping alone yields “soft safety” (avoidance behavior without guarantees), (ii) post-hoc CBF filtering prevents collisions but can cause abrupt corrections if the policy was not trained with the filter in the loop, and (iii) retraining the …
-
WhatWhen2Ask: Cost-Aware LLM Querying for Autonomous Robots in Uncertain Environments
… challenges such as partial observability, sparse rewards, and long-horizon planning. While reinforcement learning (RL) enables agents to learn from experience, standard policies often struggle to generalize in the presence of ambiguous tasks or incomplete information. Large language models (LLMs) …
-
Generative Discovery via Reinforcement Learning
… their discovery potential: Mitigate the bias of reward shaping. RL relies on reward signals from trial-anderror experience, but these signals can be sparse, meaning they are only provided once a desired solution is found and otherwise zero. Most trials, therefore, offer little to no feedback. A …
-
Achieving Robustness and Generalization in MARL for Sequential Social Dilemmas through Bilinear Value Networks
… games. Recent literature has demonstrated that reward shaping can not only be used to enable MARL agents to discover diverse, human-interpretable strategies with emergent qualities, but also help alleviate the issue in conventional actor-critic methods that tend to converge to suboptimal Nash …
-
Who Should I Trust? Uncertainty and Risk for Knowledge Transfer from Multiple Sources in Reinforcement Learning Domains
… risk-neutral transfer -- namely potential-based reward shaping and successor features -- and extend them to the risk-sensitive setting. We empirically validate all our contributions on standard RL benchmarks, where they are shown to outperform other state-of-the-art transfer learning approaches …
-
The Application of Reinforcement Learning for Interceptor Guidance
… combining the PPO algorithm with a specialized reward shaping method and tuned parameters for the engagements of interest. Low-fidelity vehicle models were used to reduce training time and narrow the scope of work towards improving the guidance algorithms. Models were trained and tested on …
-
Deep learning based approaches for imitation learning.
… and reinforcement learning. A deep reward shaping method is proposed that learns a potential reward function from demonstrations. Finally, memory architectures in deep neural networks are investigated to provide context to the agent when taking actions. Using recurrent neural …
-
Reinforcement Learning and Reward Estimation for Dialogue Policy Optimisation
… system to learn to act optimally by maximising a reward function. This reward function is designed to induce the system behaviour required for goal-oriented applications, which usually means fulfilling the user’s goal as efficiently as possible. However, in real-world spoken dialogue systems, the …
-
Cyber-Physical-Social Systems for Autonomous Defense: Enabling Mission-Centric, Adaptive, and Anytime Intelligence
… further incorporates expert feedback into reward shaping, enabling interpretable, adaptive, and trustworthy defense strategies in safety-critical vehicular environments. Task AIF introduces an Anytime Inference (AIF) algorithm for SBNs that supports incremental, resource-aware reasoning. …