Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 25 for “"Markov Decision Processes (MDPs)"”.

  1. Tensor decomposition and parallelization of Markov Decision Processes

    Markov Decision Processes (MDPs) with large state spaces arise frequently when applied to real world problems. Optimal solutions to such problems exist, but may not be computationally tractable, as the required processing scales exponentially with the number of states. Unsurprisingly, investigating …

    mit Repository record for Tensor decomposition and parallelization of Markov Decision Processes (opens in a new tab)

  2. Robust, risk-sensitive, and data-driven control of Markov Decision Processes

    Markov Decision Processes (MDPs) model problems of sequential decision-making under uncertainty. They have been studied and applied extensively. Nonetheless, there are two major barriers that still hinder the applicability of MDPs to many more practical decision making problems: * The decision

    mit Repository record for Robust, risk-sensitive, and data-driven control of Markov Decision Processes (opens in a new tab)

  3. Continuous state Q-learning

    … solution technique developed to solve classical Markov Decision Processes, MDPs. Markov Decision Processes are models for sequential decision making problems and address many classical control problems. In Chapter I, this paper discusses the model and some standard solution techniques used in …

    ttu Repository record for Continuous state Q-learning (opens in a new tab)

  4. Kernel-based approximate dynamic programming using Bellman residual elimination

    Many sequential decision-making problems related to multi-agent robotic systems can be naturally posed as Markov Decision Processes (MDPs). An important advantage of the MDP framework is the ability to utilize stochastic system models, thereby allowing the system to make sound decisions even if …

    mit Repository record for Kernel-based approximate dynamic programming using Bellman residual elimination (opens in a new tab)

  5. Planning under uncertainty with Bayesian nonparametric models

    … on Representation Expansion (SCORE) for learning Markov Decision Processes (MDPs) that exhibit an underlying multiple-model structure. SCORE addresses the co-dependence between observation clustering and model expansion. The second contribution provides a realtime, non-myopic, risk-aware planning …

    mit Repository record for Planning under uncertainty with Bayesian nonparametric models (opens in a new tab)

  6. Fast approximate hierarchical solution of MDPs

    … and solving hierarchical models of large Markov decision processes (MDPs). As the size of the MDP increases, finding an exact solution becomes intractable, so we expect only to find an approximate solution. We also assume that the hierarchies we create are not necessarily applicable to …

    mit Repository record for Fast approximate hierarchical solution of MDPs (opens in a new tab)

  7. Stochastic Dynamic Optimization and Games in Operations Management

    … and game models in operations management.</p><p>Markov decision processes (MDPs) and sequential games are good models of many real sequential decision processes. However, in diverse applications in operations research and economics, the state of the MDP is a vector and the curse of dimensionality …

    ohiolink Repository record for Stochastic Dynamic Optimization and Games in Operations Management (opens in a new tab)

  8. Towards an optimised and adaptive automation of post-incident malware investigation: a novel reinforcement learning framework

    … Framework. As a result of leveraging advanced Markov Decision Processes (MDPs), the framework enhances decision-making capabilities, enabling efficient and accurate analysis of malware in live memory dumps. This study introduces a unified investigation model that integrates static analysis, …

    london-metro Repository record for Towards an optimised and adaptive automation of post-incident malware investigation: a novel reinforcement learning framework (opens in a new tab)

  9. Bayesian Theory of Mind : modeling human reasoning about beliefs, desires, goals, and social relations

    … of approximately rational planning, such as Markov decision processes (MDPs), partially observable MDPs (POMDPs), and Markov games. ToM reasoning will be formalized as rational probabilistic inference over these models of intentional (inter)action, termed Bayesian Theory of Mind (BToM). …

    mit Repository record for Bayesian Theory of Mind : modeling human reasoning about beliefs, desires, goals, and social relations (opens in a new tab)

  10. Computationally Efficient Reinforcement Learning under Partial Observability

    … latent state of the system. Partially observable Markov decision processes (POMDPs) are a generalization of Markov decision processes (MDPs) that model this challenge. Unfortunately, planning and learning near-optimal policies in POMDPs is computationally intractable. Most existing algorithms …

    mit Repository record for Computationally Efficient Reinforcement Learning under Partial Observability (opens in a new tab)

  11. Simulating Dynamical Systems from Data

    … offers an opportunity for automated data-driven decision-making in various domains. However, a significant barrier to realizing this potential is the issues inherent to these datasets: high-dimensionality, noise, sparsity, and confounding. In this thesis, we propose methods to exploit the …

    mit Repository record for Simulating Dynamical Systems from Data (opens in a new tab)

  12. Experimental Design in Operations

    … value of operational models—particularly Markov Decision Processes (MDPs)—in experimental design. In Chapter 5, we address the challenge of estimating long-term cumulative outcomes, such as customer lifetime value, using short-term experimental data. We develop novel inference methods …

    mit Repository record for Experimental Design in Operations (opens in a new tab)

  13. Local multiagent control in large factored planning Problems

    … the solution of general, cooperative multiagent Markov Decision Processes (MDPs). To achieve this, the proposed approximation architectures assume that the solution of the overall system can be represented with sparsely interacting (i.e., local) value function components that -- if found -- …

    mit Repository record for Local multiagent control in large factored planning Problems (opens in a new tab)

  14. Data, models and decisions for large-scale stochastic optimization problems

    Modern business decisions exceed human decision making ability: often, they are of a large scale, their outcomes are uncertain, and they are made in multiple stages. At the same time, firms have increasing access to data and models. Faced with such complex decisions and increasing access to data …

    mit Repository record for Data, models and decisions for large-scale stochastic optimization problems (opens in a new tab)

  15. Uniform positive recurrence and long term behavior of Markov decision processes, with applications in sensor scheduling

    … and value iteration algorithms for discrete-time Markov decision processes (MDPs). First, we adapt two recent results in controlled diffusion processes to suit countable state MDPs by making assumptions that approximate continuous behavior. We show that if the MDP is stable under any stationary …

    texas Repository record for Uniform positive recurrence and long term behavior of Markov decision processes, with applications in sensor scheduling (opens in a new tab)

  16. Online Reinforcement Learning in Factored Markov Decision Processes and Unknown Markov Games

    … has been shown to achieve minimax optimality for Markov decision processes (MDPs) in the tabular case. However, such a model may be too general for some problems where certain structures allow for much more efficient learning. In the first part of this thesis, we consider the factored MDP model, …

    mit Repository record for Online Reinforcement Learning in Factored Markov Decision Processes and Unknown Markov Games (opens in a new tab)

  17. Statistical Methods for Off-Policy Learning

    Sequential decision problems are ubiquitous and have been studied across many areas of science, engineering, and business. The task of learning good policies from historical records collected under another policy—commonly referred to as off-policy learning—has mainly been studied in two settings: …

    cambridge Repository record for Statistical Methods for Off-Policy Learning (opens in a new tab)

  18. Design and Analysis of Defect- and Fault-tolerant Nano-Computing Systems

    … inherent variability in nanoscale fabrication processes. Compared to current CMOS devices, nanodevices are also more susceptible to signal noise and thermal perturbations. One approach for developing robust digital systems from such unreliable nanodevices is to introduce defect- and …

    vt Repository record for Design and Analysis of Defect- and Fault-tolerant Nano-Computing Systems (opens in a new tab)

  19. Dynamic Discrete Choice Estimation using Reinforcement Learning with Applications in Online Food Markets

    … models are widely used to analyze sequential decision-making in economics and marketing. However, their estimation remains computationally challenging, especially as state spaces expand, limiting their application to large-scale consumer datasets. This thesis develops Reinforcement Learning …

    cambridge Repository record for Dynamic Discrete Choice Estimation using Reinforcement Learning with Applications in Online Food Markets (opens in a new tab)

  20. Verification of linear-time properties for finite probabilistic systems

    … In this thesis we explore this question for Markov Decision Processes (MDPs), which are finite state models involving stochastic and non-deterministic behaviour over discrete time steps. The kind of specifications we focus on are those that describe the correctness of individual executions of …

    uiuc Repository record for Verification of linear-time properties for finite probabilistic systems (opens in a new tab)

Page 1 of 2