Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 44 for “"Markov decision process (MDP)"”.

  1. Cognitive Radar Applied To Target Tracking Using Markov Decision Processes

    … coexistence problem is modeled as a Markov decision process (MDP), and reinforcement learning is applied to drive the radar to optimal behavior.

    vt Repository record for Cognitive Radar Applied To Target Tracking Using Markov Decision Processes (opens in a new tab)

  2. Creation of a Cognitive Radar with Machine Learning: Simulation and Implementation

    … by modelling the radar environment as a Markov Decision Process (MDP), and then apply Deep-Q Learning to optimize radar performance. The radar environment includes a single point target and a communications system that will potentially interfere with the radar. We demonstrate that the …

    vt Repository record for Creation of a Cognitive Radar with Machine Learning: Simulation and Implementation (opens in a new tab)

  3. Inverse Reinforcement Learning and Routing Metric Discovery

    … as a reinforcement learning (RL) agent and a Markov decision process (MDP). The problem of routing metric discovery is then posed as a problem of recovering the reward function, given observed optimal behavior. We show that this approach is empirically suited for determining the relative …

    vt Repository record for Inverse Reinforcement Learning and Routing Metric Discovery (opens in a new tab)

  4. Cooperative communications in wireless networks : novel approaches in the mac layer

    … infrastructure. Secondly, we design a novel Markov decision process (MDP) framework for the cooperative retransmission problem in the wireless networks. This MDP model is proven to be simple, yet very efficient approach for distributed optimization and decision making in the cooperation …

    nus Repository record for Cooperative communications in wireless networks : novel approaches in the mac layer (opens in a new tab)

  5. Predicting Chemical Reactions at the MechanisticLevel through Deep Reinforcement Learning

    … bonds. We then define a molecular environment Markov Decision Process (MDP) that codifies the allowed mechanistic steps and evaluates them by utilizing a thermodynamic energy oracle as the reward function. To solve this environment, we build a graph neural network-based policy and value …

    mit Repository record for Predicting Chemical Reactions at the MechanisticLevel through Deep Reinforcement Learning (opens in a new tab)

  6. Reinforcement Learning, Modeling Markets, and Professional Basketball Free Agency

    … approach to modeling and optimizing decision-making in professional basketball free agency and related economic environments. A Markov Decision Process (MDP) framework is introduced to capture the strategic interactions of NBA teams bidding for free agents under budgetary and roster …

    chapman Repository record for Reinforcement Learning, Modeling Markets, and Professional Basketball Free Agency (opens in a new tab)

  7. Compact parametric models for efficient sequential decision making in high-dimensional, uncertain domains

    … a single agent can autonomously make sequential decisions in large, high-dimensional, uncertain domains. This thesis presents decision-making algorithms for maximizing the expected sum of future rewards in two types of large, high-dimensional, uncertain situations: when the agent knows its …

    mit Repository record for Compact parametric models for efficient sequential decision making in high-dimensional, uncertain domains (opens in a new tab)

  8. Detailed Inventory Record Inaccuracy Analysis

    … data sources an infinite horizon discounted Markov decision process (MDP) is generated and optimized. Moreover, the traditional cost based reward structure is abandoned to put more emphasis on the effects of IRI. Instead a new measure is developed as inventory performance by combining four …

    arkansas Repository record for Detailed Inventory Record Inaccuracy Analysis (opens in a new tab)

  9. Approximate dynamic programming with applications in multi-agent systems

    … problem. Next, we formulate this problem as a Markov Decision Process (MDP) and present a system architecture designed to improve mission-level functional reliability through system self-awareness and adaptive mission planning. Since most multi-agent mission problems are computationally …

    mit Repository record for Approximate dynamic programming with applications in multi-agent systems (opens in a new tab)

  10. Adaptive Collaborative Channel Finding Approaches for Autonomous Marine Vehicles

    … PBACS is compared to lawnmower surveying and to Markov decision process (MDP) planning with two state-of-the-art reward functions: Upper Confidence Bound (UCB) and Maximum Value Information (MVI). The performance of each method is evaluated through comparison of the time it takes to identify a …

    mit Repository record for Adaptive Collaborative Channel Finding Approaches for Autonomous Marine Vehicles (opens in a new tab)

  11. Resource allocation problems in stochastic sequential decision making

    … arise in the context of stochastic sequential decision making problems. The practical utility of optimal algorithms for these problems is limited due to their high computational and storage requirements. Also, an increasing number of applications require a decentralized solution. We develop …

    mit Repository record for Resource allocation problems in stochastic sequential decision making (opens in a new tab)

  12. Learning to Plan by Learning Rules

    … Temporal Logic (LTL) formulas in a hierarchical Markov Decision Process (MDP). We achieve satisfaction by planning over the hierarchical MDP using a modified version of value iteration. We achieve composability by building off of a hierarchical reinforcement learning (HRL) framework called the …

    mit Repository record for Learning to Plan by Learning Rules (opens in a new tab)

  13. Trucking: novel spot-market dispatching models

    … a dynamic assignment problem, implemented as a Markov Decision Process (MDP), which has its objective as maximizing the operation profit at the end of the dispatching planning horizon. A freight spot-market loads generation platform is created to mimic the dynamics of trucks and loads in such …

    wayne-thes Repository record for Trucking: novel spot-market dispatching models (opens in a new tab)

  14. Value learning through Bellman residuals and neural function approximations in deterministic systems

    … is the problem of learning values in a Markov Decision Process (MDP) by optimizing for the Bellman residuals directly without using any heuristic surrogates. Compared to the traditional Approximate Dynamic Programming (ADP) methods, this approach can have both advantages and …

    uiuc Repository record for Value learning through Bellman residuals and neural function approximations in deterministic systems (opens in a new tab)

  15. Data-driven coordination of assets in power distribution systems for ancillary service provision

    … coordination problem is cast as a multi-stage decision-making problem and formulated as a Markov decision process (MDP), in which the unknown power injections are modeled as uncertainty sources. The MDP is solved via a reinforcement learning algorithm to obtain a control policy that maps the …

    uiuc Repository record for Data-driven coordination of assets in power distribution systems for ancillary service provision (opens in a new tab)

  16. Dynamic marketing policies : constructing Markov states for reinforcement learning

    … We interpret sequential targeting problems as a Markov Decision Process (MDP), which can be solved using a range of Reinforcement Learning (RL) algorithms. MDPs require the construction of Markov state spaces. These state spaces summarize the current information about each customer in each time …

    mit Repository record for Dynamic marketing policies : constructing Markov states for reinforcement learning (opens in a new tab)

  17. Selective prioritisation of real-time IP packets to improve quality of service in 5G wireless sensor networks

    … agent. By framing the network scheduling as a Markov Decision Process (MDP), a novel Transformer-based RL agent is trained online. The findings demonstrate that this agent can learn an optimal policy in real time, achieving 100% scheduling success with ultralow latency and closing the loop from …

    london-metro Repository record for Selective prioritisation of real-time IP packets to improve quality of service in 5G wireless sensor networks (opens in a new tab)

  18. Electric Vehicle Fleet Charging Management

    … to optimize charge schedules, encompassing decisions related to timing and charging rates. To alleviate computational complexity, we propose a construction warmstart heuristic to expedite the generation of feasible solutions. The second problem focuses on charge scheduling for EV fleets at …

    alabama Repository record for Electric Vehicle Fleet Charging Management (opens in a new tab)

  19. Development of drilling optimization models for autonomous rotary drilling systems.

    … strength (UCS) and ROP, enabling optimal decision-making protocols. To evaluate optimized operating procedure, this thesis presents a comparative study of surface operating parameters using weight on bit (WOB), and rotary speed (RPM) versus drilling mechanical specific energy (DMSE), and …

    rgu Repository record for Development of drilling optimization models for autonomous rotary drilling systems. (opens in a new tab)

  20. Value function approximation architectures for neuro-dynamic programming

    … approximations of the transition kernel of the Markov decision process (MDP). Lastly, we apply the proposed results of value function approximation techniques to several applications. In the power management model, we focus on the processor speed control problem to balance the performance and …

    uiuc Repository record for Value function approximation architectures for neuro-dynamic programming (opens in a new tab)

Page 1 of 3