Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 9 of 9 for “"temporal difference learning"”.

  1. Explorations of the practical issues of learning prediction-control tasks using temporal difference learning methods

    Thesis (M.S.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 1993.

    mit Repository record for Explorations of the practical issues of learning prediction-control tasks using temporal difference learning methods (opens in a new tab)

  2. Sparse Value Function Approximation for Reinforcement Learning

    <p>A key component of many reinforcement learning (RL) algorithms is the approximation of the value function. The design and selection of features for approximation in RL is crucial, and an ongoing area of research. One approach to the problem of feature selection is to apply sparsity-inducing …

    duke Repository record for Sparse Value Function Approximation for Reinforcement Learning (opens in a new tab)

  3. Efficient reinforcement learning via singular value decomposition, end-to-end model-based methods and reward shaping

    Reinforcement learning (RL) provides a general framework for data-driven decision making. However, the very same generality that makes this approach applicable to a wide range of problems is also responsible for its well-known inefficiencies. In this thesis, we consider different properties which …

    mit Repository record for Efficient reinforcement learning via singular value decomposition, end-to-end model-based methods and reward shaping (opens in a new tab)

  4. Learning optimal discourse strategies in a spoken dialogue system

    … thesis investigates the issues involved with learning optimal discourse strategies on the basis of experience gained through conversations between human users and natural language agents. A spoken dialogue agent, ELVIS, is implemented as a testbed for learning optimal discourse strategies. …

    mit Repository record for Learning optimal discourse strategies in a spoken dialogue system (opens in a new tab)

  5. The Impact of Threat on Behavioral and Neural Markers of Learning in Anxiety

    … Decision science and in particular reinforcement learning models provide a quantitative framework to explain how the likelihood and value of such outcomes are estimated, thus allowing the measurement of parameters of decision-making that may differ between high- and low- anxiety groups. However, …

    vt Repository record for The Impact of Threat on Behavioral and Neural Markers of Learning in Anxiety (opens in a new tab)

  6. Designing policy optimization algorithms for multi-agent reinforcement learning

    Multi-agent reinforcement learning (RL) studies the sequential decision-making problem in the setting where multiple agents exist in an environment and jointly determine the environment transition. The relationship between the agents can be cooperative, competitive, or mixed depending on how the …

    gatech Repository record for Designing policy optimization algorithms for multi-agent reinforcement learning (opens in a new tab)

  7. Models of aposematism and the role of aversive learning

    … of co-evolution and the mechanisms of aversive learning are at the heart of the current research. On the one hand, to explain stability and persistence of aposematic signals requires a theory of co-evolution of defence and signals. On the other hand, the role of the predator and details of the …

    city-london Repository record for Models of aposematism and the role of aversive learning (opens in a new tab)

  8. Evolving Neural Networks with HyperNEAT and Online Training

    … through strategic application of online learning. Several methods are proposed and explored. All methodologies are tested using a team gathering task. A simulated environment is setup with gathering robots that must locate resources and work together to carry the resources back to a …

    texas-state Repository record for Evolving Neural Networks with HyperNEAT and Online Training (opens in a new tab)

  9. Hraní nedeterministických her s učením

    Práce se věnuje studiu a implementaci metod použitých pro učení z průběhu hraní. Zvolenou hrou pro tuhle práci jsou Vrhcáby. Algoritmus použitý pro učení neuronové sítě se nazývá učení z časového rozdílu s použitím stop vhodnosti. Tento algoritmus je známý i pod jménem TD(lambda). V teoretické …

    brno-tech Repository record for Hraní nedeterministických her s učením (opens in a new tab)