Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 32 for “"value iteration"”.

  1. Solving Large MDPs Quickly with Partitioned Value Iteration

    Value iteration is not typically considered a viable algorithm for solving large-scale MDPs because it converges too slowly. However, its performance can be dramatically improved by eliminating redundant or useless backups, and by backing up states in the right order. We present several methods …

    byu Repository record for Solving Large MDPs Quickly with Partitioned Value Iteration (opens in a new tab)

  2. Approximate value iteration approaches to constrained dynamic portfolio problems

    … suboptimal solution methods based on approximate value iteration. The primary innovation is the use of mean-variance portfolio selection methods. We present two case studies that employ these approximate value iteration methods. The first case study explores the effect of an insolvency constraint …

    mit Repository record for Approximate value iteration approaches to constrained dynamic portfolio problems (opens in a new tab)

  3. Improved worst-case regret bounds for randomized least-squares value iteration

    … work studies regret minimization with randomized value functions in reinforcement learning. In tabular finite-horizon Markov Decision Processes, we introduce a clipping variant of one classical Thompson Sampling (TS)-like algorithm, randomized least-squares value iteration (RLSVI). Our …

    uiuc Repository record for Improved worst-case regret bounds for randomized least-squares value iteration (opens in a new tab)

  4. Uniform positive recurrence and long term behavior of Markov decision processes, with applications in sensor scheduling

    … relating to stability, optimal control, and value iteration algorithms for discrete-time Markov decision processes (MDPs). First, we adapt two recent results in controlled diffusion processes to suit countable state MDPs by making assumptions that approximate continuous behavior. We show that …

    texas Repository record for Uniform positive recurrence and long term behavior of Markov decision processes, with applications in sensor scheduling (opens in a new tab)

  5. Acceleration of Iterative Methods for Markov Decision Processes

    … widely employed to solve MDP problems include value iteration and policy iteration. Although simple to implement, these approaches are, nevertheless, limited in the size of problems that can be solved, due to excessive computation required to find close-to-optimal solutions. My thesis proposes …

    toronto-retro Repository record for Acceleration of Iterative Methods for Markov Decision Processes (opens in a new tab)

  6. A tensor-train-decomposition-based algorithm for high-dimensional pursuit-evasion games

    … Best-Response Tensor-Train-decomposition-based Value Iteration (BR-TT-VI) was developed. BR-TT-VI builds on concepts from game theory, dynamic programming (DP), and tensor-train decomposition. By using TT decomposition, BR-TT-VI greatly reduces the effects of the curse of dimensionality. This …

    mit Repository record for A tensor-train-decomposition-based algorithm for high-dimensional pursuit-evasion games (opens in a new tab)

  7. Approximate Dynamic Programming with Applications

    … control problem is characterized by the optimal value function. For a large class of problems the optimal value function must satisfy a Hamilton-Jacobi-Bellman type equation. Two common methods for solving such equations are policy iteration and value iteration. Both these methods are studied in …

    lund Repository record for Approximate Dynamic Programming with Applications (opens in a new tab)

  8. Localisation and navigation in GPS-denied environments using RFID tags

    … tags are used for localisation, subsequently value iteration is used to navigate to a defined goal. Results are presented, concluding that it is feasible to localise and navigate using only RFID tags, in simulation. Localisation feasibility is also confirmed by experimental measurements.

    cape-town Repository record for Localisation and navigation in GPS-denied environments using RFID tags (opens in a new tab)

  9. Applicability of deep learning approaches to non-convex optimization for trajectory-based policy search

    … against globally optimal policies determined via value iteration on simple control tasks. Second, three systems built for parallelized, non-convex optimization across trajectories with a shared neural network constraint are described and analyzed. Finally, techniques from deep learning known to …

    mit Repository record for Applicability of deep learning approaches to non-convex optimization for trajectory-based policy search (opens in a new tab)

  10. Towards practical neural network meta-modeling

    … architecture using Q-learning, a popular value iteration algorithm from the reinforcement learning community for sequential decision problems. On the task of object classification, the Q-learning agent outperforms all human crafted models that are similar to those in the search space. By …

    mit Repository record for Towards practical neural network meta-modeling (opens in a new tab)

  11. Multi-Player Zero-Sum Markov Games with Networked Separable Interactions

    … a Markov non-stationary NE and provide finite-iteration guarantees for a series of value-iteration-based algorithms. We also provide numerical experiments to corroborate our theoretical results.

    mit Repository record for Multi-Player Zero-Sum Markov Games with Networked Separable Interactions (opens in a new tab)

  12. Real-time maneuvering decisions for autonomous air combat

    … decisions for autonomous one-on-one air combat. Value iteration techniques are used to compute a function approximation representing the solution to the dynamic program. The function approximation is then used as a maneuvering policy for UAS autonomous air combat. The result is an algorithm …

    mit Repository record for Real-time maneuvering decisions for autonomous air combat (opens in a new tab)

  13. Solving Cyber-Alert Allocation Markov Games with Deep Reinforcement Learning

    … via the use of dynamic programming and Q-maximin value iteration based algorithms. We then move into approximation techniques, using deep neural networks and Q-learning to derive near-optimal strategies that allow us to explore much larger models. We assess the effectiveness of our allocation …

    texas-state Repository record for Solving Cyber-Alert Allocation Markov Games with Deep Reinforcement Learning (opens in a new tab)

  14. Risk-Aware Neural Navigation for Interactive Driving

    … risk associated with nearby agents, (2) value iteration using the risk map to learn a policy, and (3) a Trajectory Sampler, which samples from this policy to generate a trajectory. We uniquely evaluate our planner in an interactive manner, adjusting the surroundings at each time step, and …

    mit Repository record for Risk-Aware Neural Navigation for Interactive Driving (opens in a new tab)

  15. Discrete Approximate Information States in Partially Observable Environments

    … can solve for the globally optimal policy using value iteration for the DAIS model, allowing us to disambiguate the performance of the AIS objective from the policy search. Going further, for small problems with finite information states, we reformulate the DAIS learning problem as a novel …

    mit Repository record for Discrete Approximate Information States in Partially Observable Environments (opens in a new tab)

  16. Private and Provably Efficient Federated Decision-Making

    … We propose variants of least-squares value iteration algorithms that are provably no-regret with only a constant communication budget. We believe that the future of machine learning entails large-scale cooperation between various data-driven entities, and this work will be beneficial …

    mit Repository record for Private and Provably Efficient Federated Decision-Making (opens in a new tab)

  17. Learning to Plan by Learning Rules

    … the hierarchical MDP using a modified version of value iteration. We achieve composability by building off of a hierarchical reinforcement learning (HRL) framework called the options framework, in which low-level options can be composed arbitrarily. And lastly, we achieve data-efficient learning …

    mit Repository record for Learning to Plan by Learning Rules (opens in a new tab)

  18. Dynamic Programming and Time-Varying Delay Systems

    … of solving the inequality in the paper is value iteration, which is shown to work well in many different applications. In the second part of the thesis, two analysis methods for systems with time-varying delays are presented in two papers. The first paper presents a set of simple graphical …

    lund Repository record for Dynamic Programming and Time-Varying Delay Systems (opens in a new tab)

  19. On the design and implementation of decision-theoretic, interactive, and vision-driven mobile robots

    … the computational efficiency of the point-based value iteration algorithm while tackling the problem of multi-step actions using Dynamic Bayesian Networks. In addition, we describe a state-of-the-art simultaneous localization and mapping algorithm for robots equipped with stereo vision. We first …

    ubc Repository record for On the design and implementation of decision-theoretic, interactive, and vision-driven mobile robots (opens in a new tab)

  20. Optimization for stochastic, partially observed systems using a sampling-based approach to learn switched policies

    … interest. Our approach is based on a method like value iteration to learn a switching law. Because the POMDP problem is intractable, we use a Monte Carlo approximation to evaluate system behavior and optimize a switching law based on sampling. We explicitly analyze the sensitivity of expected cost …

    uiuc Repository record for Optimization for stochastic, partially observed systems using a sampling-based approach to learn switched policies (opens in a new tab)

Page 1 of 2