Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 48 for “"Actor-critic"”.
-
Actor-critic algorithms
… policies. In this thesis, we propose and study actor-critic algorithms which combine the above two approaches with simulation to find the best policy among a parameterized class of policies. Actor-critic algorithms have two learning units: an actor and a critic. An actor is a decision maker with …
-
Reinforcement Actor-Critic Learning As A Rehearsal In MicroRTS
… it for a different RL framework, called actor-critic RL. We show that on the one hand the actor-critic framework allows RLaR to be much simpler, but on the other hand it leaves room for a key component of RLaR--a prediction function that relates a learner's observations with that of its …
-
Reinforcement learning for power scheduling in a grid-tied pv-battery electric vehicles charging station
… Q-learning algorithm and an advantage actor-critic (A2C) algorithm, in performing power scheduling in the EV charging station under static conditions. To assess the performances of the proposed algorithms, the conventional Q-learning and actor-critic algorithm were implemented to …
-
Deep Reinforcement Learning for Robotic Tasks: Manipulation and Sensor Odometry
… uses HER and many independent instances of actors and critics from the DDPG to increase a robot's learning effectiveness. AACHER is used to evaluate the results in both custom and existing robot environments.In the first part of our research, we discuss the LIMO algorithm, an odometry …
-
Learning in the Multi-Robot Pursuit Evasion Game
… second one is a modified version of the fuzzy-actor critic learning (FACL) algorithm, which is called fuzzy actor-critic learning Automaton (FACLA) algorithm. It uses the continuous actor-critic learning Automaton (CACLA) algorithm to tune the parameters of the FIS.After that, a decentralized …
-
Learning-based Decision Making in Wireless Communications
… optimization methods to quickly solve critical decision-making problems. With this motivation, in this thesis, machine learning methods are developed and utilized for obtaining optimal/near-optimal solutions for timely decision making in wireless networks.</p><p>Content caching at the …
-
Experience-driven Control for Networking and Computing
… two new techniques, TE-aware exploration and actor-critic-based prioritized experience replay, to optimize the general DRL framework particularly for TE. Furthermore, we propose an Actor-Critic-based Transfer learning framework for TE, ACT-TE, which solves a practical problem in …
-
Experience-driven Control For Networking And Computing
… two new techniques, TE-aware exploration and actor-critic-based prioritized experience replay, to optimize the general DRL framework particularly for TE. Furthermore, we propose an Actor-Critic-based Transfer learning framework for TE, ACT-TE, which solves a practical problem in …
-
Artificial Neural Network-Based Robotic Control
… policy gradients (DDPG) algorithm, an actor-critic reinforcement learning strategy, originally conceived by Google DeepMind. After training, the robot performs controlled locomotion within an enclosed area. The paper also details the robot design process and explores the challenges of …
-
Winning at Pokémon Random Battles Using Reinforcement Learning
… employs a Monte Carlo Tree Search informed by a actor-critic network trained using Proximal Policy Optimization with experience collected through self-play. The agent peaked at rank 8 (1693 Elo) on the official Pokémon Showdown gen4randombattles ladder, which is the best known performance by any …
-
Reinforcement learning in network control
… learning methods such as Q-Learning, Actor-Critic, etc. are heuristic and do not offer performance guarantees. In contrast, model-based learning methods offer performance guarantees, but can only be applied with bounded state spaces. In the thesis, we propose to use model-based …
-
Reinforcement Learning and Reward Estimation for Dialogue Policy Optimisation
… Two sample-efficient algorithms, trust region actor-critic with experience replay (TRACER) and episodic natural actor-critic with experience replay (eNACER), are introduced. In addition, a corpus of demonstration data is utilised to pre-train the models prior to on-line reinforcement learning …
-
Market making in dry waters : reinforcement learning strategies for market making in illiquid markets
… RL algorithms: Deep Q-Networks (DQN), Advantage Actor-Critic (A2C) and Proximal Policy Optimization (PPO). They are evaluated in a simulated stock market environment on their performance in liquid and illiquid market conditions. Findings show that DQN outperforms the other in both conditions. The …
-
Modeling Naturalistic Driver Behavior in Traffic Using Machine Learning
… during car-following situations and safety critical events. Driving behavior is considered as a human decision process in this research which provides opportunities for an artificial driver agent simulator to learn according to naturalistic driving data. This thesis presents two mechine …
-
A Comprehensive Study on Energy Management Strategy Design of an Extended-Range Electric Vehicle
… The comparative results show that the Soft Actor-Critic (SAC) had a 36% faster convergence speed than a traditional algorithm while providing a smoother and more stable action space. The fuel consumption with SAC also outplays by around 3%, which achieves almost 90% of the DP results.
-
DEEP REINFORCEMENT LEARNING FOR BUILDING ENERGY MANAGEMENT
… Deep Reinforce- ment Learning algorithm with an actor-critic framework and Trust Region Policy, to control a thermal energy storage scheduling problem in a continuous state-action space with stochastic electric generation. To this end, two main steps are carried out in this thesis. First, the PPO …
-
Zero-shot learning to execute tasks with robots
… Second, we explore the use of more powerful actor-critic methods, augmented with hindsight experience replay (HER). We determine that approaches requiring low-dimensional representations of the environment, such as HER, will not scale gracefully to handle more complex environments. Finally, …
-
Performance of an AGI-aspiring system & narrow-AI approaches : a systematic comparison
… tested on a double deep Q network (DDQ) and an actor-critic (AC). ONA was found to vastly outperform the narrow-AI learners in proficiency and learning speed, but could not handle one variable being random or its actions being swapped. The results indicate that the DDQ is the most general of the …
-
Evolutionary algorithms for optimising reinforcement learning policy approximation
… In particular, the A3C (asynchronous advantage actor critic) approach demonstrated in Mnih et al. (2016) was able to halve the training time of the existing state-of-the-art approaches. However, these methods still require relatively large amounts of training resources due to the fundamental …
-
Applied optimal control for dynamically stable legged locomotion
… on our real robot. I describe, in detail, the actor-critic reinforcement learning algorithm that is implemented on the return map dynamics of the biped. Finally, I address issues of scaling and controller augmentation using tools from optimal control theory and a simulation of a planar one-leg …
Page 1 of 3