Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 44 for “"Policy Optimization"”.
-
Firewall Policy Optimization and Management
Firewalls enforce a security policy by inspecting packets arriving or departing a network. This is accomplished by sequentially comparing the policy rules with the header of an arriving packet until the first match is found. This process becomes time consuming as policies become larger and more …
-
Preferential proximal policy optimization in reinforcement learning
The Proximal Policy Optimization (PPO), a policy gradient method, excels in reinforcement learning with its ”surrogate” objective function and stochastic gradient ascent. However, PPO does not fully consider the significance of frequently encountered states in policy/value updates. To address this, …
-
Designing policy optimization algorithms for multi-agent reinforcement learning
… algorithms essentially solve a bi-level optimization problem by tracking an artificial auxiliary variable in addition to the decision variable and updating them at different rates. We propose a two-time-scale stochastic gradient descent method under a special type of gradient oracle which …
-
Learning-based optimal and robust control: A policy optimization perspective
… robotic tasks. Central to this success is policy optimization (PO), a subclass of RL where policies—parameterized mappings from observations to actions—are iteratively optimized to enhance system performance. PO’s flexibility, scalability, and data-driven nature make it effective for …
-
Low-thrust Spacecraft guidance and control using proximal policy optimization
… Reinforcement Learning (RL) algorithm Proximal Policy Optimization (PPO) for solving low-thrust spacecraft guidance and control problems. First, an agent is trained to complete a 302 day mass-optimal low-thrust transfer between the Earth and Mars. This is accomplished while only providing the …
-
Simulation-Based Reinforcement Learning Policy Optimization for Tactile Manipulation: A Case Study on the Eyesight Hand
… formulations, and curriculum strategies on policy performance. Our findings highlight the benefits of delta position control, a carefully selected observation space including joint states, control vectors, object pose, and contact forces, and success-driven curriculum learning. Our study …
-
Machine learning for high performance computing applications
… and reinforcement learning using proximal policy optimization. Included are three different advancements applying these techniques. The first application used K-means clustering and Gradient Boosted Tree Regression (GBTR) to predict estimated queue time for jobs submitted to an HPC system. …
-
Reinforcement Learning for Autonomous Aircraft Control and Aerial Maneuvering in Simulated Environments
… in a Unity-based simulation. Using Proximal Policy Optimization (PPO), the agent learns to fly, avoid obstacles, and reach targets through continuous control. A combination of environment design and reward shaping was essential for achieving stable learning. Results show that performance …
-
Winning at Pokémon Random Battles Using Reinforcement Learning
… by a actor-critic network trained using Proximal Policy Optimization with experience collected through self-play. The agent peaked at rank 8 (1693 Elo) on the official Pokémon Showdown gen4randombattles ladder, which is the best known performance by any non-human agent for this format. This strong …
-
Learning and decentralized control in linear switched systems
… the second part of this dissertation, we turn to policy gradient approaches with the goal of starting from some sub-optimal controller and using measurement data to optimize the control gains. Recently, policy optimization for control purposes has received renewed attention due to the increasing …
-
Selecting appropriate reinforcement-learning algorithms for robot manipulation domains
… Evolution Strategy, Deep Deterministic Policy Gradients, and Proximal Policy Optimization. We compare their performance on various target domains to measure quantitatively their dependence on varied features of the environment. We study the effect of actuation noise, observation noise, …
-
Multiple-part-type systems in high volume manufacturing : Kanban System design for automatic production scheduling
… on the parameter settings of the Control-Point Policy, the optimum Kanban levels are obtained. The simulation software Simul8 was used to model the factory line and the Kanban system. Using the optimum Kanban levels, the Kanban system will act as an automatic production scheduling system that …
-
Reinforcement Learning–Based Discrete Prompt Optimization for Neuro-Symbolic Structured Simplification of Complex Game Descriptions with Large Language Models
… formalizes simplification as a discrete prompt optimization problem and introduces a neuro-symbolic pipeline that maps raw natural language into controlled GameChangineer sentences via scenario normalization, retrieval-augmented code generation, and AST-based FACTS extraction. A reinforcement …
-
Market making in dry waters : reinforcement learning strategies for market making in illiquid markets
… (DQN), Advantage Actor-Critic (A2C) and Proximal Policy Optimization (PPO). They are evaluated in a simulated stock market environment on their performance in liquid and illiquid market conditions. Findings show that DQN outperforms the other in both conditions. The research highlights the …
-
Hybrid AI-driven Approach to Context-Aware Inter-Slice Load Balancing for Cloud-Native Functions in 5G Networks
… models like deep Q-networks (DQN) and proximal policy optimization (PPO) for traffic load balancing across core network slices, enhancing service quality. Kubernetes-integrated deployments use Prometheus for real-time predictive analytics, expanding beyond network-centric metrics to …
-
LightMARL : smart swarm coordination in urban spaces
… This thesis introduces LightMARL, an optimization framework that facilitates effective multi-agent coordination on resource-constrained edge devices while ensuring coordination quality and real-time performance. LightMARL tackles three primary challenges: computational efficiency, …
-
Predicting Chemical Reactions at the MechanisticLevel through Deep Reinforcement Learning
… we build a graph neural network-based policy and value network, where the policy network is first pre-trained on an open database of elementary radical reactions (RMechDB). Then, we use proximal policy optimization (PPO) to fine-tune the model and predict reasonable reaction mechanisms …
-
Reinforcement Learning, Modeling Markets, and Professional Basketball Free Agency
… environment is developed, leveraging Proximal Policy Optimization (PPO) to approximate optimal policies for team decision-making.</p> <p>Empirical results demonstrate that the RL agent successfully learns strategic bidding behavior that aligns with dynamic programming benchmarks in simplified …
-
Adversarial Prompt Transformation for Systematic Jailbreaks of LLMs
… into a successful jailbreak. Thus it learns a policy based on relation to existing jailbreak prompts that informs the generator LLM of what makes an adversarial prompt successful. This was implemented using Proximal Policy Optimization (PPO) and tested with both a classifier and judge reward …
-
ACADIA: Efficient and Robust Adversarial Attacks Against Deep Reinforcement Learning
… Deep-Q Learning Network (DQN) and Proximal Policy Optimization (PPO), under Atari games and MuJoCo where both targeted and non-targeted attacks are considered with or without the state-of-the-art defenses in DRL (i.e., RADIAL and ATLA). Our results demonstrate that the proposed ACADIA …
Page 1 of 3