Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 8 of 8 for “"Policy Gradient Methods"”.
-
Sample Complexity of Incremental Policy Gradient Methods for Solving Multi-Task Reinforcement Learning
… this problem, we are interested in studying the gradient approach, which iteratively updates an estimate of the optimal policy using the gradients of the value functions. The classic policy gradient method, however, may be expensive to implement in the multi-task settings as it requires access to …
-
An Introduction to Reinforcement Learning
… bandits, fitted dynamic programming algorithms, policy gradient methods, imitation learning, and tree search-based planning methods. Our contribution to the RL literature is an approachable and concise presentation of core RL algorithms that balances practical considerations with theoretical …
-
Reinforcement learning for multi-agent and robust control systems
… In contrast to existing solvers, we introduce policy-gradient methods to solve the robust control problem, with global convergence guarantees, despite its nonconvexity. More interestingly, we show that two of these methods enjoy the implicit regularization property: the iterates of the …
-
Scalable hierarchical evolution strategies
… performance to be comparable to state-of-the art policy gradient methods. However, S-ES has not been tested in conjunction with HRL methods, which empower temporal abstraction thus allowing agents to tackle more challenging problems. We introduce a novel method merging S-ES and HRL, which creates …
-
Learning and decentralized control in linear switched systems
… first part of the thesis, we develop synthesis methods for decentralized control of switched systems with mode-dependent (more generally, path-dependent) performance specifications. This specification flexibility is important when achievable system performance varies greatly between modes, as a …
-
Preferential proximal policy optimization in reinforcement learning
The Proximal Policy Optimization (PPO), a policy gradient method, excels in reinforcement learning with its ”surrogate” objective function and stochastic gradient ascent. However, PPO does not fully consider the significance of frequently encountered states in policy/value updates. To address this, …
-
Maximum entropy on-policy reinforcement learning with monotonic policy improvement
This Thesis was approved for publication on 2023-07-21 at 16:57.
-
Multi-agent reinforcement learning: A mean-field perspective
Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-03-28 without embargo terms