Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 17 of 17 for “"policy iteration"”.

  1. An approach for nonlinear control design via approximate dynamic programming

    … using dynamic programming algorithms such as policy iteration. Exact policy iteration is computationally infeasible for systems of even moderate dimension, which leads us to consider methods based on Approximate Policy Iteration. In such methods, we first select an approximation architecture …

    mit Repository record for An approach for nonlinear control design via approximate dynamic programming (opens in a new tab)

  2. Acceleration of Iterative Methods for Markov Decision Processes

    … employed to solve MDP problems include value iteration and policy iteration. Although simple to implement, these approaches are, nevertheless, limited in the size of problems that can be solved, due to excessive computation required to find close-to-optimal solutions. My thesis proposes a new …

    toronto-retro Repository record for Acceleration of Iterative Methods for Markov Decision Processes (opens in a new tab)

  3. Trustworthy Reinforcement Learning under Constraints and Perturbations

    … (DBCE) and the Density-Based Correlated Policy Iteration (DBCPI) algorithm. (2) To address sequential resource allocation with situational constraints, we develop a primal–dual Situational Constraint RL (SCRL) framework, introducing a density-based formulation to measure resource …

    auckland-ms Repository record for Trustworthy Reinforcement Learning under Constraints and Perturbations (opens in a new tab)

  4. Approximate Dynamic Programming with Applications

    … common methods for solving such equations are policy iteration and value iteration. Both these methods are studied in this thesis.An approximate policy iteration algorithm is presented for both the continuous and discrete time settings. It is shown that the sequence produced by this algorithm …

    lund Repository record for Approximate Dynamic Programming with Applications (opens in a new tab)

  5. Learning-based Optimal Control of Time-Varying Linear Systems Over Large Time Intervals

    … This allows the utilization of an off-policy iteration method to learn the controller gains. We show that the performance of the learning-based controller approximates that of the model-based optimal controller and the approximation accuracy improves as the control problem’s time …

    vt Repository record for Learning-based Optimal Control of Time-Varying Linear Systems Over Large Time Intervals (opens in a new tab)

  6. Resource allocation and load-shedding policies based on Markov decision processes for renewable energy generation and storage

    … utilizes a Markov Decision Process with backward policy iteration. This is based on a probabilistic method that chooses the best load-shedding path that minimizes the expected total cost to ensure no power failure. We compare our results with two control policies, a load-balancing policy and a …

    ucf

  7. A constrained MDP-based vertical handoff decision algorithm for wireless networks

    … when making the handoff decisions. The policy iteration and Q-learning algorithms are employed to determine the optimal policy. Structural results on the optimal vertical handoff policy are derived by using the concept of supermodularity. We show that the optimal policy is a threshold …

    ubc Repository record for A constrained MDP-based vertical handoff decision algorithm for wireless networks (opens in a new tab)

  8. Playing Tetris with deep reinforcement learning

    … algorithm based on evolution algorithms and policy iteration. It uses manually crafted features and achieves 51 million lines cleared in an average game. In recent years, deep reinforcement learning (DRL) has achieved outstanding performance with Atari and Go games. An initial attempt by …

    uiuc Repository record for Playing Tetris with deep reinforcement learning (opens in a new tab)

  9. A new approach to multistage serial inventory systems

    … that over a finite horizon an echelon basestock policy is optimal. Federgruen and Zipkin (1984) extend their result to the infinite-horizon case for both discounted and average costs. We present a new approach to this multistage serial inventory management problem, and give new proofs of these …

    mit Repository record for A new approach to multistage serial inventory systems (opens in a new tab)

  10. Adaptive Fuzzy Reinforcement Learning for Flock Motion Control

    … on an online fuzzy reinforcement learning Value Iteration scheme which is precise and flexible. This distributed adaptive control system simultaneously targets a number of flocking objectives; namely: 1) tracking the leader, 2) keeping a safe distance from the neighboring agents, and 3) reaching …

    ottawa-retro Repository record for Adaptive Fuzzy Reinforcement Learning for Flock Motion Control (opens in a new tab)

  11. Agile load transportation systems using aerial robots

    … In addition, we use an online least square policy iteration algorithm. At the end, we propose a high level algorithm for navigation in cluttered environments considering a quadrotor with suspended load. Furthermore, distributed control of multiple quadrotors with suspended load is addressed …

    unm Repository record for Agile load transportation systems using aerial robots (opens in a new tab)

  12. Continuous low-rank tensor decompositions, with applications to stochastic optimal control and data assimilation

    … Next, we develop compressed versions of value iteration, policy iteration, and multilevel algorithms for solving dynamic programming problems arising in stochastic optimal control. These techniques enable computing global solutions to a broader set of problems, for example those with non-affine …

    mit Repository record for Continuous low-rank tensor decompositions, with applications to stochastic optimal control and data assimilation (opens in a new tab)

  13. Sparse Value Function Approximation for Reinforcement Learning

    … second contribution extends LARS-TD to integrate policy optimization with sparse value learning. We extend the <italic>L<sub>1</sub></italic> regularized linear fixed point to include a maximum over policies, defining a new, "greedy" fixed point. The greedy fixed point adds a new invariant to the …

    duke Repository record for Sparse Value Function Approximation for Reinforcement Learning (opens in a new tab)

  14. Stochastic models for analysis and optimization of unmanned aerial vehicle delivery on last-mile logistics

    … in the problem and determine the optimal policy for decision-makers by applying a policy iteration algorithm. To overcome of computational challenges, a novel approximation method called the decomposition-based approach is proposed to split the original Markov decision problem for the …

    ksu Repository record for Stochastic models for analysis and optimization of unmanned aerial vehicle delivery on last-mile logistics (opens in a new tab)

  15. Learning from expert advice framework: Algorithms and applications

    … as a Markov decision process (MDP) and apply policy iteration to solve it. For the logarithmic loss, we prove that the optimal strategy for the adversary is the greedy policy, whereas for the absolute loss, in the $2$-experts, discounted cost setting, we prove that the optimal strategy is a …

    uiuc Repository record for Learning from expert advice framework: Algorithms and applications (opens in a new tab)

  16. On the resolution of misspecification in stochastic optimization, variational inequality, and game-theoretic problems

    … (1) We first show that a misspecified value iteration scheme converges almost surely to its true counterpart and the mean-squared error after $K$ iterations is $O(\sqrt{1/K})$; (2) An analogous asymptotic almost-sure convergence statement is provided for misspecified policy iteration; and (3) …

    uiuc Repository record for On the resolution of misspecification in stochastic optimization, variational inequality, and game-theoretic problems (opens in a new tab)

  17. Fast numerical algorithms for optimal robot motion planning

    Optimization of high-level autonomous tasks requires solving the optimal motion planning problem for a mobile robot. For example, to reach the desired destination on time, a self-driving car must quickly navigate streets and avoid hazardous obstacles such as buildings or other cars, as well as …

    uiuc Repository record for Fast numerical algorithms for optimal robot motion planning (opens in a new tab)