Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 46 for “"policy gradient"”.
-
Retrospective Policy Gradient
Policy Gradient (PG) methods represent a notable category of reinforcement learning algorithms, particularly effective in continuous action domains. These techniques adjust the parameters of parametric policies through stochastic gradient ascent, typically utilizing on-policy trajectory samples to …
-
Imitation Learning with Superhuman Policy Gradient Optimization for Sequential Cancer Treatment Decisions
… cancer (HNC) treatment. Our method, Superhu- man Policy Gradient Optimization (SPGO), integrates inverse reinforcement learning principles with policy gradient updates to derive three-stage treat- ment policies directly from recorded physician decisions. A pre-trained clinical simulator—combining …
-
Sample Complexity of Incremental Policy Gradient Methods for Solving Multi-Task Reinforcement Learning
… this problem, we are interested in studying the gradient approach, which iteratively updates an estimate of the optimal policy using the gradients of the value functions. The classic policy gradient method, however, may be expensive to implement in the multi-task settings as it requires access to …
-
Learning and decentralized control in linear switched systems
… the second part of this dissertation, we turn to policy gradient approaches with the goal of starting from some sub-optimal controller and using measurement data to optimize the control gains. Recently, policy optimization for control purposes has received renewed attention due to the increasing …
-
Sample-Efficient Deep Reinforcement Learning for Continuous Control
… for continuous action problems; Interpolated Policy Gradient (IPG), unifying prior policy gradient algorithm variants through theoretical analyses on bias and variance; and Temporal Difference Models (TDM), interpreting a parameterized Q-function as a generalized dynamics model for novel …
-
An Introduction to Reinforcement Learning
… bandits, fitted dynamic programming algorithms, policy gradient methods, imitation learning, and tree search-based planning methods. Our contribution to the RL literature is an approachable and concise presentation of core RL algorithms that balances practical considerations with theoretical …
-
Investigating Reinforcement Learning and Evolutionary Computation for Games with Stochasticity and Incomplete Information
… two deep reinforcement learning methods: policy gradient and evolutionary strategies for training the neural network behind the AI players for Ticket to Ride, a complex strategic board game. By comparing AI players’ performance and policies with existing heuristics players, we show that …
-
Motor learning on a heaving plate via improved-SNR algorithms
… motivated by a novel Signal-to-Noise Ratio for policy gradient algorithms, are developed, and shown to provide more efficient learning in noisy environments. These algorithms are then demonstrated on a heaving foil, where it is shown to learn a flapping gait on an experimental system orders of …
-
Reinforcement learning for multi-agent and robust control systems
… In contrast to existing solvers, we introduce policy-gradient methods to solve the robust control problem, with global convergence guarantees, despite its nonconvexity. More interestingly, we show that two of these methods enjoy the implicit regularization property: the iterates of the …
-
Preferential proximal policy optimization in reinforcement learning
The Proximal Policy Optimization (PPO), a policy gradient method, excels in reinforcement learning with its ”surrogate” objective function and stochastic gradient ascent. However, PPO does not fully consider the significance of frequently encountered states in policy/value updates. To address this, …
-
Consistent Depth Estimation in Data-Driven Simulation for Autonomous Driving
… practicality. We train several end-to-end policy gradient models in varying versions of VISTA, each utilizing a different depth method, and see that end-to-end models trained in the consistent depth version of VISTA deviate least from the human driven center line.
-
Reinforcement Learning Control for Mobile Robot Parking with Safety Constraints
… obstacle avoidance. We apply Deep Deterministic Policy Gradient (DDPG) methods for continuous control and evaluate policies across three Simulink environments of increasing fidelity: a kinematic model, a dynamic model, and a dynamic model with actuator disturbance. In parking tasks, DDPG learns …
-
Scalable hierarchical evolution strategies
… performance to be comparable to state-of-the art policy gradient methods. However, S-ES has not been tested in conjunction with HRL methods, which empower temporal abstraction thus allowing agents to tackle more challenging problems. We introduce a novel method merging S-ES and HRL, which creates …
-
A Neuro-Symbolic Reinforcement Learning Architecture: Integrating Perception, Reasoning, and Control
… presented. In this environment, an analysis of policy-gradient-based reinforcement learning algorithms is given. Then, by leveraging the performance of deep learning with the semantic reasoning and interpretability of symbolically defined program- ming, a novel neuro-symbolic learning method is …
-
Geometry of Feedback Control and Learning
… directly, viewing control synthesis by policy gradient based algorithms. Adopting such a point of view has been partially inspired by the success of learning algorithms, such as Reinforcement Learning (RL), where using principles of Dynamic Programming (DP), one can devise real-time …
-
Development and Deployment of a Dynamic Soaring Capable UAV using Reinforcement Learning
… by flying through regions of vertical wind gradient such as the wind shear layer. With reinforcement learning (RL), a fixed wing unmanned aerial vehicle (UAV) can be trained to perform DS maneuvers optimally for a variety of wind shear conditions. To accomplish this task, a 6-degreesof- …
-
Cognitive GPR for subsurface sensing based on edge computing and deep reinforcement learning
… learning method called deep deterministic policy gradient (DDPG) with a new reward function derived from 3D GPR data. The proposed methods are evaluated using GPR modeling and simulation software called GprMax. Simulation results show that our proposed cognitive GPRs outperform other GPR …
-
Towards realising multimodal robots
… is proposed. ReCoAl works by allowing a direct policy gradient based Reinforcement Learning algorithm to improve the controller of every evolved robot to better utilise the available morphological resources before the fitness evaluation. The findings indicate that the learning process has both …
-
Designing Intelligent Energy Efficient Scheduling Algorithm To Support Massive IoT Communication In LoRa Networks
… scheduling algorithm, a Deep Deterministic policy gradient algorithm with channel activity detection (CAD) to optimize the energy efficiency of LoRaWAN in cross-layer architecture in massive IoT with star topology. We also design a CAD-based simulator for evaluating any algorithms with …
-
Improving cache replacement policy using deep reinforcement learning
… quickly and efficiently. A cache's replacement policy plays a major role in determining the cache's effectiveness and performance. The replacement policy is an algorithm that chooses which piece of data in the cache should be evicted when the cache becomes full and new elements are requested. In …
Page 1 of 3