Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 106 for “"Markov Decision Processes"”.

  1. Hazard avoidance alerting with Markov decision processes

    (cont.) (incident rate and unnecessary alert rate), the MDP-based logic can meet or exceed that of alternate logics.

    mit Repository record for Hazard avoidance alerting with Markov decision processes (opens in a new tab)

  2. Simulation-based optimization of Markov decision processes

    Markov decision processes have been a popular paradigm for sequential decision making under uncertainty. Dynamic programming provides a framework for studying such problems, as well as for devising algorithms to compute an optimal control policy. Dynamic programming methods rely on a suitably …

    mit Repository record for Simulation-based optimization of Markov decision processes (opens in a new tab)

  3. Acceleration of Iterative Methods for Markov Decision Processes

    This research focuses on Markov Decision Processes (MDP). MDP is one of the most important and challenging areas of Operations Research. Every day people make many decisions: today's decisions impact tomorrow's and tomorrow's will impact the ones made the day after. Problems in Engineering, …

    toronto-retro Repository record for Acceleration of Iterative Methods for Markov Decision Processes (opens in a new tab)

  4. Tensor decomposition and parallelization of Markov Decision Processes

    Markov Decision Processes (MDPs) with large state spaces arise frequently when applied to real world problems. Optimal solutions to such problems exist, but may not be computationally tractable, as the required processing scales exponentially with the number of states. Unsurprisingly, investigating …

    mit Repository record for Tensor decomposition and parallelization of Markov Decision Processes (opens in a new tab)

  5. Learning bounded optimal behavior using Markov decision processes

    … metalevel optimization problem is posed within a Markov Decision Process framework and is solved off-line to determine a policy for carrying out computations. Once the optimal policy is determined, it serves efficiently as an online metalevel controller that selects computational actions …

    mit Repository record for Learning bounded optimal behavior using Markov decision processes (opens in a new tab)

  6. Cognitive Radar Applied To Target Tracking Using Markov Decision Processes

    … coexistence problem is modeled as a Markov decision process (MDP), and reinforcement learning is applied to drive the radar to optimal behavior.

    vt Repository record for Cognitive Radar Applied To Target Tracking Using Markov Decision Processes (opens in a new tab)

  7. Model-free reinforcement learning in non-stationary Markov Decision Processes

    … unknown environment, usually modeled by a Markov Decision Process (MDP). The classical RL literature typically assumes that the state transition functions and the reward functions of the MDP are time-invariant. Such a stationary model, however, cannot capture the dynamic nature of many …

    uiuc Repository record for Model-free reinforcement learning in non-stationary Markov Decision Processes (opens in a new tab)

  8. Robust, risk-sensitive, and data-driven control of Markov Decision Processes

    Markov Decision Processes (MDPs) model problems of sequential decision-making under uncertainty. They have been studied and applied extensively. Nonetheless, there are two major barriers that still hinder the applicability of MDPs to many more practical decision making problems: * The decision

    mit Repository record for Robust, risk-sensitive, and data-driven control of Markov Decision Processes (opens in a new tab)

  9. Policy-based average-reward and robust Markov decision processes and reinforcement learning

    Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms

    uiuc Repository record for Policy-based average-reward and robust Markov decision processes and reinforcement learning (opens in a new tab)

  10. Online Reinforcement Learning in Factored Markov Decision Processes and Unknown Markov Games

    … has been shown to achieve minimax optimality for Markov decision processes (MDPs) in the tabular case. However, such a model may be too general for some problems where certain structures allow for much more efficient learning. In the first part of this thesis, we consider the factored MDP model, …

    mit Repository record for Online Reinforcement Learning in Factored Markov Decision Processes and Unknown Markov Games (opens in a new tab)

  11. Approximate solution methods for partially observable Markov and semi-Markov decision processes

    … infinite-horizon partially observable Markov and semi-Markov decision processes (POMDP and POSMDP). One of the main contributions of this thesis is a lower cost approximation method for finite-space POMDPs with the average cost criterion, and its extensions to semi-Markov partially …

    mit Repository record for Approximate solution methods for partially observable Markov and semi-Markov decision processes (opens in a new tab)

  12. Dealing with uncertainty : a comparison of robust optimization and partially observable Markov decision processes

    … it. This thesis is intended to be used by a decision maker to determine how best to formulate a problem. Robust optimization and partially observable Markov decision processes (POMDPs) are two methods of dealing with uncertainty in real life problems. Robust optimization is used primarily in …

    mit Repository record for Dealing with uncertainty : a comparison of robust optimization and partially observable Markov decision processes (opens in a new tab)

  13. Hierarchical decomposition of multi-agent Markov decision processes with application to health aware planning

    … approach (HD-MMDP) for solving Multi-agent Markov Decision Processes (MMDPs), which is a natural framework for formulating stochastic sequential decision-making problems. In particular, the HD-MMDP algorithm builds a decomposition structure by exploiting coupling relationships in the reward …

    mit Repository record for Hierarchical decomposition of multi-agent Markov decision processes with application to health aware planning (opens in a new tab)

  14. Practical reinforcement learning using representation learning and safe exploration for large scale Markov decision processes

    … agents who can solve stochastic sequential decision making problems through interacting with the environment is the promise of Reinforcement Learning (RL), scaling existing RL methods to realistic domains such as planning for multiple unmanned aerial vehicles (UAVs) has remained a challenge …

    mit Repository record for Practical reinforcement learning using representation learning and safe exploration for large scale Markov decision processes (opens in a new tab)

  15. Uniform positive recurrence and long term behavior of Markov decision processes, with applications in sensor scheduling

    … and value iteration algorithms for discrete-time Markov decision processes (MDPs). First, we adapt two recent results in controlled diffusion processes to suit countable state MDPs by making assumptions that approximate continuous behavior. We show that if the MDP is stable under any stationary …

    texas Repository record for Uniform positive recurrence and long term behavior of Markov decision processes, with applications in sensor scheduling (opens in a new tab)

  16. Resource allocation and load-shedding policies based on Markov decision processes for renewable energy generation and storage

    … method discussed in this thesis utilizes a Markov Decision Process with backward policy iteration. This is based on a probabilistic method that chooses the best load-shedding path that minimizes the expected total cost to ensure no power failure. We compare our results with two control …

    ucf

  17. Decentralized control of multi-robot systems using partially observable Markov Decision Processes and belief space macro-actions

    … Decentralized Partially Observable Markov Decision Processes (Dec-POMDPs) are general models for multi-robot coordination problems. However, representing and solving Dec-POMDPs is often intractable for large problems. This thesis extends the Dec-POMDP framework to the Decentralized …

    mit Repository record for Decentralized control of multi-robot systems using partially observable Markov Decision Processes and belief space macro-actions (opens in a new tab)

  18. Continuous state Q-learning

    … solution technique developed to solve classical Markov Decision Processes, MDPs. Markov Decision Processes are models for sequential decision making problems and address many classical control problems. In Chapter I, this paper discusses the model and some standard solution techniques used in …

    ttu Repository record for Continuous state Q-learning (opens in a new tab)

  19. Efficient Bayesian Nonparametric Methods for Model-Free Reinforcement Learning in Centralized and Decentralized Sequential Environments

    … RL in centralized and decentralized sequential decision-making problems using BNPMs. We show how the control policies can be learned efficiently under model-free RL schemes with BNPMs. Specifically, for centralized sequential decision-making, we study Q learning with Gaussian processes to solve …

    duke Repository record for Efficient Bayesian Nonparametric Methods for Model-Free Reinforcement Learning in Centralized and Decentralized Sequential Environments (opens in a new tab)

  20. Improved worst-case regret bounds for randomized least-squares value iteration

    … learning. In tabular finite-horizon Markov Decision Processes, we introduce a clipping variant of one classical Thompson Sampling (TS)-like algorithm, randomized least-squares value iteration (RLSVI). Our $\tilde{\mathrm{O}}(H^2S\sqrt{AT})$ high-probability worst-case regret bound …

    uiuc Repository record for Improved worst-case regret bounds for randomized least-squares value iteration (opens in a new tab)

Page 1 of 6