Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 157 for “"markov decision process"”.

  1. Optimizing credit limit policy by Markov Decision Process Models

    … in recent years. Lenders realize their operation decisions are crucial in determining how much profit is achieved from a card. This thesis focuses on the most well-known operating policy: the management of credit limit. Lenders traditionally applied static decision models to manage the credit …

    soton Repository record for Optimizing credit limit policy by Markov Decision Process Models (opens in a new tab)

  2. An extended Kalman filter extension of the augmented Markov decision process

    … planners, such as the partially observable Markov decision process, plan over the entire set of beliefs (distributions over the robot's position). Unfortunately, this approach is only tractable for problems with very few states. Between these two extreme approaches, however, lies a continuum …

    mit Repository record for An extended Kalman filter extension of the augmented Markov decision process (opens in a new tab)

  3. Manufacturing Technology Adoption in Dynamic Product Environments

    … of the technology investment problem: Stochastic Processes in the form of a Partially Observable Markov Decision Process, Non-Linear Mathematical Programming, and Computer Simulation integrated with a Multi-attribute Model based on measurement theory.

    uiuc Repository record for Manufacturing Technology Adoption in Dynamic Product Environments (opens in a new tab)

  4. Probabilistic roadmaps in uncertain environments

    … of some edges are unknown. This is modelled as a decision-theoretic planning problem described through a partially observable Markov Decision Process (POMDP). It is shown that the optimal policy can depend on accounting for the value of information from observations. The model scalability and the …

    birmingham Repository record for Probabilistic roadmaps in uncertain environments (opens in a new tab)

  5. Data-Driven sequential decision making with learning under ambiguity

    Markov decision processes are often used to model sequential decision-making problems in uncertain dynamic environments, such as equipment maintenance and replacement problems, and inventory control problems. The objective of these problems is to find a policy or strategy, which is a prescription …

    ttu Repository record for Data-Driven sequential decision making with learning under ambiguity (opens in a new tab)

  6. Dynamic resource allocation in CDMA cellular communications systems

    … formulate the resource allocation problem as a Markov decision process. Due to the enormous size of the state space, applying the traditional solution technique, dynamic programming, is impractical. We therefore consider approximation techniques. As a first step towards simplification, we divide …

    mit Repository record for Dynamic resource allocation in CDMA cellular communications systems (opens in a new tab)

  7. Cognitive Radar Applied To Target Tracking Using Markov Decision Processes

    … coexistence problem is modeled as a Markov decision process (MDP), and reinforcement learning is applied to drive the radar to optimal behavior.

    vt Repository record for Cognitive Radar Applied To Target Tracking Using Markov Decision Processes (opens in a new tab)

  8. Solving Dec-MDPs with options and intention recognition

    … We model the problems with a Decentralized Markov Decision Process, and we make use of options and intention recognition to solve the problem. Rather than directly solving the Dec-MDP, which is NEXP-Complete, we instead solve a set of single-agent MDPs, that we can solve in P-Complete, and …

    mit Repository record for Solving Dec-MDPs with options and intention recognition (opens in a new tab)

  9. Information-theoretic Algorithms for Model-free Reinforcement Learning

    … algorithm for infinte-horizon, average-reward decision processes where the transition function has a finite yet unknown dependence on history, and where the induced Markov Decision Process is assumed to be weakly communicating. This algorithm combines the Lempel-Ziv (LZ) parsing tree structure …

    mit Repository record for Information-theoretic Algorithms for Model-free Reinforcement Learning (opens in a new tab)

  10. Creation of a Cognitive Radar with Machine Learning: Simulation and Implementation

    … by modelling the radar environment as a Markov Decision Process (MDP), and then apply Deep-Q Learning to optimize radar performance. The radar environment includes a single point target and a communications system that will potentially interfere with the radar. We demonstrate that the …

    vt Repository record for Creation of a Cognitive Radar with Machine Learning: Simulation and Implementation (opens in a new tab)

  11. Symbolic planning in belief space

    … for a satisfying plan to a partially observable Markov decision process, or a POMDP, while benefiting from advantages of classical symbolic planning such as compact belief state expression, domain-independent heuristics, and structural simplicity. Belief space symbolic formalism, an extension of …

    mit Repository record for Symbolic planning in belief space (opens in a new tab)

  12. Inverse Reinforcement Learning and Routing Metric Discovery

    … as a reinforcement learning (RL) agent and a Markov decision process (MDP). The problem of routing metric discovery is then posed as a problem of recovering the reward function, given observed optimal behavior. We show that this approach is empirically suited for determining the relative …

    vt Repository record for Inverse Reinforcement Learning and Routing Metric Discovery (opens in a new tab)

  13. An integrated performance model learning and planning approach for optimal infrastructure facility maintenance under partial observability

    Infrastructure inspection and maintenance decision making is a stochastic and partially observable problem. This thesis presents a learning and decision-making approach for developing optimal joint inspection and maintenance policies for civil infrastructure facilities under performance model …

    tdl Repository record for An integrated performance model learning and planning approach for optimal infrastructure facility maintenance under partial observability (opens in a new tab)

  14. Sampling-based algorithms for stochastic optimal control

    … propose a novel algorithm called the incremental Markov Decision Process (iMDP) to compute incrementally any-time control policies that approximate arbitrarily well an optimal policy in terms of the expected cost. The main idea is to generate a sequence of finite discretizations of the original …

    mit Repository record for Sampling-based algorithms for stochastic optimal control (opens in a new tab)

  15. An analytical model of MAC protocol dependant power consumption in multi-hop ad hoc wireless sensor networks

    … protocol. The model is formulated as a semi-Markov decision process (SMDP) wherein the node's states, sojourn times, and transition probabilities are controlled by a virtual node controller. The overall operation of a communication protocol is viewed as a randomized policy for the SMDP, and …

    njit Repository record for An analytical model of MAC protocol dependant power consumption in multi-hop ad hoc wireless sensor networks (opens in a new tab)

  16. Resource allocation and load-shedding policies based on Markov decision processes for renewable energy generation and storage

    … method discussed in this thesis utilizes a Markov Decision Process with backward policy iteration. This is based on a probabilistic method that chooses the best load-shedding path that minimizes the expected total cost to ensure no power failure. We compare our results with two control …

    ucf

  17. Cooperative communications in wireless networks : novel approaches in the mac layer

    … infrastructure. Secondly, we design a novel Markov decision process (MDP) framework for the cooperative retransmission problem in the wireless networks. This MDP model is proven to be simple, yet very efficient approach for distributed optimization and decision making in the cooperation …

    nus Repository record for Cooperative communications in wireless networks : novel approaches in the mac layer (opens in a new tab)

  18. TOWARDS HUMAN-CENTRIC AI: INVERSE REINFORCEMENT LEARNING MEETS ALGORITHMIC FAIRNESS

    … translates various fairness principles into fair decisions by translating them to a combination of utility, non-causal and causal components which are, in turn, mapped to elements of a Constrained Markov Decision Process. Our final work unifies the concepts of value alignment and fairness by …

    nus Repository record for TOWARDS HUMAN-CENTRIC AI: INVERSE REINFORCEMENT LEARNING MEETS ALGORITHMIC FAIRNESS (opens in a new tab)

  19. Planning under uncertainty for dynamic collision avoidance

    … the problem within the Partially Observable Markov Decision Process (POMDP) framework, and use generic MDP/POMDP solvers offline to compute vertical-only avoidance strategies that optimize a cost function to balance flight-plan deviation with risk of collision. We then describe a second …

    mit Repository record for Planning under uncertainty for dynamic collision avoidance (opens in a new tab)

Page 1 of 8