Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 25 for “"Proximal Policy Optimization"”.

  1. Preferential proximal policy optimization in reinforcement learning

    The Proximal Policy Optimization (PPO), a policy gradient method, excels in reinforcement learning with its ”surrogate” objective function and stochastic gradient ascent. However, PPO does not fully consider the significance of frequently encountered states in policy/value updates. To address this, …

    uoit Repository record for Preferential proximal policy optimization in reinforcement learning (opens in a new tab)

  2. Low-thrust Spacecraft guidance and control using proximal policy optimization

    … use of the Reinforcement Learning (RL) algorithm Proximal Policy Optimization (PPO) for solving low-thrust spacecraft guidance and control problems. First, an agent is trained to complete a 302 day mass-optimal low-thrust transfer between the Earth and Mars. This is accomplished while only …

    mit Repository record for Low-thrust Spacecraft guidance and control using proximal policy optimization (opens in a new tab)

  3. Machine learning for high performance computing applications

    … system, and reinforcement learning using proximal policy optimization. Included are three different advancements applying these techniques. The first application used K-means clustering and Gradient Boosted Tree Regression (GBTR) to predict estimated queue time for jobs submitted to an HPC …

    ksu Repository record for Machine learning for high performance computing applications (opens in a new tab)

  4. Reinforcement Learning for Autonomous Aircraft Control and Aerial Maneuvering in Simulated Environments

    … navigation in a Unity-based simulation. Using Proximal Policy Optimization (PPO), the agent learns to fly, avoid obstacles, and reach targets through continuous control. A combination of environment design and reward shaping was essential for achieving stable learning. Results show that …

    debrecen Repository record for Reinforcement Learning for Autonomous Aircraft Control and Aerial Maneuvering in Simulated Environments (opens in a new tab)

  5. Winning at Pokémon Random Battles Using Reinforcement Learning

    … informed by a actor-critic network trained using Proximal Policy Optimization with experience collected through self-play. The agent peaked at rank 8 (1693 Elo) on the official Pokémon Showdown gen4randombattles ladder, which is the best known performance by any non-human agent for this format. …

    mit Repository record for Winning at Pokémon Random Battles Using Reinforcement Learning (opens in a new tab)

  6. Selecting appropriate reinforcement-learning algorithms for robot manipulation domains

    … Evolution Strategy, Deep Deterministic Policy Gradients, and Proximal Policy Optimization. We compare their performance on various target domains to measure quantitatively their dependence on varied features of the environment. We study the effect of actuation noise, observation noise, …

    mit Repository record for Selecting appropriate reinforcement-learning algorithms for robot manipulation domains (opens in a new tab)

  7. Reinforcement Learning–Based Discrete Prompt Optimization for Neuro-Symbolic Structured Simplification of Complex Game Descriptions with Large Language Models

    … formalizes simplification as a discrete prompt optimization problem and introduces a neuro-symbolic pipeline that maps raw natural language into controlled GameChangineer sentences via scenario normalization, retrieval-augmented code generation, and AST-based FACTS extraction. A reinforcement …

    vt Repository record for Reinforcement Learning–Based Discrete Prompt Optimization for Neuro-Symbolic Structured Simplification of Complex Game Descriptions with Large Language Models (opens in a new tab)

  8. Market making in dry waters : reinforcement learning strategies for market making in illiquid markets

    … (DQN), Advantage Actor-Critic (A2C) and Proximal Policy Optimization (PPO). They are evaluated in a simulated stock market environment on their performance in liquid and illiquid market conditions. Findings show that DQN outperforms the other in both conditions. The research highlights …

    reykjavik Repository record for Market making in dry waters : reinforcement learning strategies for market making in illiquid markets (opens in a new tab)

  9. Hybrid AI-driven Approach to Context-Aware Inter-Slice Load Balancing for Cloud-Native Functions in 5G Networks

    … (MADRL) models like deep Q-networks (DQN) and proximal policy optimization (PPO) for traffic load balancing across core network slices, enhancing service quality. Kubernetes-integrated deployments use Prometheus for real-time predictive analytics, expanding beyond network-centric metrics to …

    carleton Repository record for Hybrid AI-driven Approach to Context-Aware Inter-Slice Load Balancing for Cloud-Native Functions in 5G Networks (opens in a new tab)

  10. Predicting Chemical Reactions at the MechanisticLevel through Deep Reinforcement Learning

    … we build a graph neural network-based policy and value network, where the policy network is first pre-trained on an open database of elementary radical reactions (RMechDB). Then, we use proximal policy optimization (PPO) to fine-tune the model and predict reasonable reaction mechanisms …

    mit Repository record for Predicting Chemical Reactions at the MechanisticLevel through Deep Reinforcement Learning (opens in a new tab)

  11. Reinforcement Learning, Modeling Markets, and Professional Basketball Free Agency

    … (RL) environment is developed, leveraging Proximal Policy Optimization (PPO) to approximate optimal policies for team decision-making.</p> <p>Empirical results demonstrate that the RL agent successfully learns strategic bidding behavior that aligns with dynamic programming benchmarks in …

    chapman Repository record for Reinforcement Learning, Modeling Markets, and Professional Basketball Free Agency (opens in a new tab)

  12. Adversarial Prompt Transformation for Systematic Jailbreaks of LLMs

    … into a successful jailbreak. Thus it learns a policy based on relation to existing jailbreak prompts that informs the generator LLM of what makes an adversarial prompt successful. This was implemented using Proximal Policy Optimization (PPO) and tested with both a classifier and judge reward …

    mit Repository record for Adversarial Prompt Transformation for Systematic Jailbreaks of LLMs (opens in a new tab)

  13. ACADIA: Efficient and Robust Adversarial Attacks Against Deep Reinforcement Learning

    … algorithms, Deep-Q Learning Network (DQN) and Proximal Policy Optimization (PPO), under Atari games and MuJoCo where both targeted and non-targeted attacks are considered with or without the state-of-the-art defenses in DRL (i.e., RADIAL and ATLA). Our results demonstrate that the proposed …

    vt Repository record for ACADIA: Efficient and Robust Adversarial Attacks Against Deep Reinforcement Learning (opens in a new tab)

  14. Computational Simulation and Machine Learning for Quality Improvement in Composites Assembly

    … interact with novel implementations of modified proximal policy optimization, based on a newly developed reinforcement learning algorithm. The resulting reinforcement learning agents are able to successfully address the underlying optimization problems that underpin the process and quality …

    vt Repository record for Computational Simulation and Machine Learning for Quality Improvement in Composites Assembly (opens in a new tab)

  15. SPIRAL: Iterative Subgraph Expansion for Knowledge-Graph Based Retrieval-Augmented Generation

    … previous work in its use of a trained, iterative policy network built on top of a prior over triples, delivering improved performance on multi-hop question answering tasks. Stage 1 trains a single-label GLASS-GNN on shortest-path heuristics, producing frozen, question-aware node embeddings at …

    mit Repository record for SPIRAL: Iterative Subgraph Expansion for Knowledge-Graph Based Retrieval-Augmented Generation (opens in a new tab)

  16. Deep Reinforcement Learning for Multirotor Flight Control: A Comparative Study of Sim-to-Real Training and Real-World Performance

    … on factors that most influence sim-to-real policy transfer. Using the Proximal Policy Optimization (PPO) algorithm, extensive ablation studies evaluated the effects of domain randomization, reward weighting, observation representation, and neural network architecture. Policies were trained …

    vt Repository record for Deep Reinforcement Learning for Multirotor Flight Control: A Comparative Study of Sim-to-Real Training and Real-World Performance (opens in a new tab)

  17. Engineering Design Automation via Imitation Learning and Reinforcement Learning

    … for surrogate design tasks like aircraft design optimization. Additionally, we assess the performance of RL agents, specifically Proximal Policy Optimization and Advantage Actor-Critic ,in the same design task. Both RL approaches achieved higher Q-scores (up to 0.99) but incurred significant …

    calgary Repository record for Engineering Design Automation via Imitation Learning and Reinforcement Learning (opens in a new tab)

  18. The Application of Reinforcement Learning for Interceptor Guidance

    … upon traditional guidance methods. Specifically, proximal policy optimization (PPO) was selected as the reinforcement learning algorithm due to its advanced and efficient nature, as well as its successful use in related work. A framework was developed and tuned for the interceptor guidance …

    vt Repository record for The Application of Reinforcement Learning for Interceptor Guidance (opens in a new tab)

  19. Orbital Maneuvers and Interplanetary Trajectory Design via Reinforcement Learning

    … of reinforcement learning (RL) to the design and optimization of low-thrust spacecraft trajectories, with an emphasis on autonomy, adaptability, and robustness in the presence of system uncertainties and unmodeled perturbations. Classical approaches to low-thrust trajectory design are …

    embry-riddle Repository record for Orbital Maneuvers and Interplanetary Trajectory Design via Reinforcement Learning (opens in a new tab)

  20. Assessment of Reinforcement Learning Algorithms for Nuclear Power Plant Fuel Optimization

    The nuclear fuel loading pattern optimization problem belongs to the class of large-scale combinatorial optimization and has been studied since the dawn of the commercial nuclear energy industry. It is also characterized by multiple objectives and constraints, which makes it impossible to solve …

    mit Repository record for Assessment of Reinforcement Learning Algorithms for Nuclear Power Plant Fuel Optimization (opens in a new tab)

Page 1 of 2