Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 32 for “"value iteration"”.
-
Solving Large MDPs Quickly with Partitioned Value Iteration
Value iteration is not typically considered a viable algorithm for solving large-scale MDPs because it converges too slowly. However, its performance can be dramatically improved by eliminating redundant or useless backups, and by backing up states in the right order. We present several methods …
-
Approximate value iteration approaches to constrained dynamic portfolio problems
… suboptimal solution methods based on approximate value iteration. The primary innovation is the use of mean-variance portfolio selection methods. We present two case studies that employ these approximate value iteration methods. The first case study explores the effect of an insolvency constraint …
-
Improved worst-case regret bounds for randomized least-squares value iteration
… work studies regret minimization with randomized value functions in reinforcement learning. In tabular finite-horizon Markov Decision Processes, we introduce a clipping variant of one classical Thompson Sampling (TS)-like algorithm, randomized least-squares value iteration (RLSVI). Our …
-
Uniform positive recurrence and long term behavior of Markov decision processes, with applications in sensor scheduling
… relating to stability, optimal control, and value iteration algorithms for discrete-time Markov decision processes (MDPs). First, we adapt two recent results in controlled diffusion processes to suit countable state MDPs by making assumptions that approximate continuous behavior. We show that …
-
Acceleration of Iterative Methods for Markov Decision Processes
… widely employed to solve MDP problems include value iteration and policy iteration. Although simple to implement, these approaches are, nevertheless, limited in the size of problems that can be solved, due to excessive computation required to find close-to-optimal solutions. My thesis proposes …
-
A tensor-train-decomposition-based algorithm for high-dimensional pursuit-evasion games
… Best-Response Tensor-Train-decomposition-based Value Iteration (BR-TT-VI) was developed. BR-TT-VI builds on concepts from game theory, dynamic programming (DP), and tensor-train decomposition. By using TT decomposition, BR-TT-VI greatly reduces the effects of the curse of dimensionality. This …
-
Approximate Dynamic Programming with Applications
… control problem is characterized by the optimal value function. For a large class of problems the optimal value function must satisfy a Hamilton-Jacobi-Bellman type equation. Two common methods for solving such equations are policy iteration and value iteration. Both these methods are studied in …
-
Localisation and navigation in GPS-denied environments using RFID tags
… tags are used for localisation, subsequently value iteration is used to navigate to a defined goal. Results are presented, concluding that it is feasible to localise and navigate using only RFID tags, in simulation. Localisation feasibility is also confirmed by experimental measurements.
-
Applicability of deep learning approaches to non-convex optimization for trajectory-based policy search
… against globally optimal policies determined via value iteration on simple control tasks. Second, three systems built for parallelized, non-convex optimization across trajectories with a shared neural network constraint are described and analyzed. Finally, techniques from deep learning known to …
-
Towards practical neural network meta-modeling
… architecture using Q-learning, a popular value iteration algorithm from the reinforcement learning community for sequential decision problems. On the task of object classification, the Q-learning agent outperforms all human crafted models that are similar to those in the search space. By …
-
Multi-Player Zero-Sum Markov Games with Networked Separable Interactions
… a Markov non-stationary NE and provide finite-iteration guarantees for a series of value-iteration-based algorithms. We also provide numerical experiments to corroborate our theoretical results.
-
Real-time maneuvering decisions for autonomous air combat
… decisions for autonomous one-on-one air combat. Value iteration techniques are used to compute a function approximation representing the solution to the dynamic program. The function approximation is then used as a maneuvering policy for UAS autonomous air combat. The result is an algorithm …
-
Solving Cyber-Alert Allocation Markov Games with Deep Reinforcement Learning
… via the use of dynamic programming and Q-maximin value iteration based algorithms. We then move into approximation techniques, using deep neural networks and Q-learning to derive near-optimal strategies that allow us to explore much larger models. We assess the effectiveness of our allocation …
-
Risk-Aware Neural Navigation for Interactive Driving
… risk associated with nearby agents, (2) value iteration using the risk map to learn a policy, and (3) a Trajectory Sampler, which samples from this policy to generate a trajectory. We uniquely evaluate our planner in an interactive manner, adjusting the surroundings at each time step, and …
-
Discrete Approximate Information States in Partially Observable Environments
… can solve for the globally optimal policy using value iteration for the DAIS model, allowing us to disambiguate the performance of the AIS objective from the policy search. Going further, for small problems with finite information states, we reformulate the DAIS learning problem as a novel …
-
Private and Provably Efficient Federated Decision-Making
… We propose variants of least-squares value iteration algorithms that are provably no-regret with only a constant communication budget. We believe that the future of machine learning entails large-scale cooperation between various data-driven entities, and this work will be beneficial …
-
Learning to Plan by Learning Rules
… the hierarchical MDP using a modified version of value iteration. We achieve composability by building off of a hierarchical reinforcement learning (HRL) framework called the options framework, in which low-level options can be composed arbitrarily. And lastly, we achieve data-efficient learning …
-
Dynamic Programming and Time-Varying Delay Systems
… of solving the inequality in the paper is value iteration, which is shown to work well in many different applications. In the second part of the thesis, two analysis methods for systems with time-varying delays are presented in two papers. The first paper presents a set of simple graphical …
-
On the design and implementation of decision-theoretic, interactive, and vision-driven mobile robots
… the computational efficiency of the point-based value iteration algorithm while tackling the problem of multi-step actions using Dynamic Bayesian Networks. In addition, we describe a state-of-the-art simultaneous localization and mapping algorithm for robots equipped with stereo vision. We first …
-
Optimization for stochastic, partially observed systems using a sampling-based approach to learn switched policies
… interest. Our approach is based on a method like value iteration to learn a switching law. Because the POMDP problem is intractable, we use a Monte Carlo approximation to evaluate system behavior and optimize a switching law based on sampling. We explicitly analyze the sensitivity of expected cost …
Page 1 of 2