Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 48 for “"Actor-critic"”.

  1. Actor-critic algorithms

    … policies. In this thesis, we propose and study actor-critic algorithms which combine the above two approaches with simulation to find the best policy among a parameterized class of policies. Actor-critic algorithms have two learning units: an actor and a critic. An actor is a decision maker with …

    mit Repository record for Actor-critic algorithms (opens in a new tab)

  2. Reinforcement Actor-Critic Learning As A Rehearsal In MicroRTS

    … it for a different RL framework, called actor-critic RL. We show that on the one hand the actor-critic framework allows RLaR to be much simpler, but on the other hand it leaves room for a key component of RLaR--a prediction function that relates a learner's observations with that of its …

    usm Repository record for Reinforcement Actor-Critic Learning As A Rehearsal In MicroRTS (opens in a new tab)

  3. Reinforcement learning for power scheduling in a grid-tied pv-battery electric vehicles charging station

    … Q-learning algorithm and an advantage actor-critic (A2C) algorithm, in performing power scheduling in the EV charging station under static conditions. To assess the performances of the proposed algorithms, the conventional Q-learning and actor-critic algorithm were implemented to …

    cape-town Repository record for Reinforcement learning for power scheduling in a grid-tied pv-battery electric vehicles charging station (opens in a new tab)

  4. Deep Reinforcement Learning for Robotic Tasks: Manipulation and Sensor Odometry

    … uses HER and many independent instances of actors and critics from the DDPG to increase a robot's learning effectiveness. AACHER is used to evaluate the results in both custom and existing robot environments.In the first part of our research, we discuss the LIMO algorithm, an odometry …

    unr Repository record for Deep Reinforcement Learning for Robotic Tasks: Manipulation and Sensor Odometry (opens in a new tab)

  5. Learning in the Multi-Robot Pursuit Evasion Game

    … second one is a modified version of the fuzzy-actor critic learning (FACL) algorithm, which is called fuzzy actor-critic learning Automaton (FACLA) algorithm. It uses the continuous actor-critic learning Automaton (CACLA) algorithm to tune the parameters of the FIS.After that, a decentralized …

    carleton Repository record for Learning in the Multi-Robot Pursuit Evasion Game (opens in a new tab)

  6. Learning-based Decision Making in Wireless Communications

    … optimization methods to quickly solve critical decision-making problems. With this motivation, in this thesis, machine learning methods are developed and utilized for obtaining optimal/near-optimal solutions for timely decision making in wireless networks.</p><p>Content caching at the …

    syracuse-diss Repository record for Learning-based Decision Making in Wireless Communications (opens in a new tab)

  7. Experience-driven Control for Networking and Computing

    … two new techniques, TE-aware exploration and actor-critic-based prioritized experience replay, to optimize the general DRL framework particularly for TE. Furthermore, we propose an Actor-Critic-based Transfer learning framework for TE, ACT-TE, which solves a practical problem in …

    syracuse-diss Repository record for Experience-driven Control for Networking and Computing (opens in a new tab)

  8. Experience-driven Control For Networking And Computing

    … two new techniques, TE-aware exploration and actor-critic-based prioritized experience replay, to optimize the general DRL framework particularly for TE. Furthermore, we propose an Actor-Critic-based Transfer learning framework for TE, ACT-TE, which solves a practical problem in …

    syracuse-diss Repository record for Experience-driven Control For Networking And Computing (opens in a new tab)

  9. Artificial Neural Network-Based Robotic Control

    … policy gradients (DDPG) algorithm, an actor-critic reinforcement learning strategy, originally conceived by Google DeepMind. After training, the robot performs controlled locomotion within an enclosed area. The paper also details the robot design process and explores the challenges of …

    calpoly Repository record for Artificial Neural Network-Based Robotic Control (opens in a new tab)

  10. Winning at Pokémon Random Battles Using Reinforcement Learning

    … employs a Monte Carlo Tree Search informed by a actor-critic network trained using Proximal Policy Optimization with experience collected through self-play. The agent peaked at rank 8 (1693 Elo) on the official Pokémon Showdown gen4randombattles ladder, which is the best known performance by any …

    mit Repository record for Winning at Pokémon Random Battles Using Reinforcement Learning (opens in a new tab)

  11. Reinforcement learning in network control

    … learning methods such as Q-Learning, Actor-Critic, etc. are heuristic and do not offer performance guarantees. In contrast, model-based learning methods offer performance guarantees, but can only be applied with bounded state spaces. In the thesis, we propose to use model-based …

    mit Repository record for Reinforcement learning in network control (opens in a new tab)

  12. Reinforcement Learning and Reward Estimation for Dialogue Policy Optimisation

    … Two sample-efficient algorithms, trust region actor-critic with experience replay (TRACER) and episodic natural actor-critic with experience replay (eNACER), are introduced. In addition, a corpus of demonstration data is utilised to pre-train the models prior to on-line reinforcement learning …

    cambridge Repository record for Reinforcement Learning and Reward Estimation for Dialogue Policy Optimisation (opens in a new tab)

  13. Market making in dry waters : reinforcement learning strategies for market making in illiquid markets

    … RL algorithms: Deep Q-Networks (DQN), Advantage Actor-Critic (A2C) and Proximal Policy Optimization (PPO). They are evaluated in a simulated stock market environment on their performance in liquid and illiquid market conditions. Findings show that DQN outperforms the other in both conditions. The …

    reykjavik Repository record for Market making in dry waters : reinforcement learning strategies for market making in illiquid markets (opens in a new tab)

  14. Modeling Naturalistic Driver Behavior in Traffic Using Machine Learning

    … during car-following situations and safety critical events. Driving behavior is considered as a human decision process in this research which provides opportunities for an artificial driver agent simulator to learn according to naturalistic driving data. This thesis presents two mechine …

    vt Repository record for Modeling Naturalistic Driver Behavior in Traffic Using Machine Learning (opens in a new tab)

  15. A Comprehensive Study on Energy Management Strategy Design of an Extended-Range Electric Vehicle

    … The comparative results show that the Soft Actor-Critic (SAC) had a 36% faster convergence speed than a traditional algorithm while providing a smoother and more stable action space. The fuel consumption with SAC also outplays by around 3%, which achieves almost 90% of the DP results.

    uts Repository record for A Comprehensive Study on Energy Management Strategy Design of an Extended-Range Electric Vehicle (opens in a new tab)

  16. DEEP REINFORCEMENT LEARNING FOR BUILDING ENERGY MANAGEMENT

    … Deep Reinforce- ment Learning algorithm with an actor-critic framework and Trust Region Policy, to control a thermal energy storage scheduling problem in a continuous state-action space with stochastic electric generation. To this end, two main steps are carried out in this thesis. First, the PPO …

    nus Repository record for DEEP REINFORCEMENT LEARNING FOR BUILDING ENERGY MANAGEMENT (opens in a new tab)

  17. Zero-shot learning to execute tasks with robots

    … Second, we explore the use of more powerful actor-critic methods, augmented with hindsight experience replay (HER). We determine that approaches requiring low-dimensional representations of the environment, such as HER, will not scale gracefully to handle more complex environments. Finally, …

    mit Repository record for Zero-shot learning to execute tasks with robots (opens in a new tab)

  18. Performance of an AGI-aspiring system & narrow-AI approaches : a systematic comparison

    … tested on a double deep Q network (DDQ) and an actor-critic (AC). ONA was found to vastly outperform the narrow-AI learners in proficiency and learning speed, but could not handle one variable being random or its actions being swapped. The results indicate that the DDQ is the most general of the …

    reykjavik Repository record for Performance of an AGI-aspiring system & narrow-AI approaches : a systematic comparison (opens in a new tab)

  19. Evolutionary algorithms for optimising reinforcement learning policy approximation

    … In particular, the A3C (asynchronous advantage actor critic) approach demonstrated in Mnih et al. (2016) was able to halve the training time of the existing state-of-the-art approaches. However, these methods still require relatively large amounts of training resources due to the fundamental …

    cape-town Repository record for Evolutionary algorithms for optimising reinforcement learning policy approximation (opens in a new tab)

  20. Applied optimal control for dynamically stable legged locomotion

    … on our real robot. I describe, in detail, the actor-critic reinforcement learning algorithm that is implemented on the return map dynamics of the biped. Finally, I address issues of scaling and controller augmentation using tools from optimal control theory and a simulation of a planar one-leg …

    mit Repository record for Applied optimal control for dynamically stable legged locomotion (opens in a new tab)

Page 1 of 3