Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 25 for “"bandit problem"”.

  1. A Control Theoretic Approach to the Stochastic Multi-armed Bandit Problem With Applications in Hyperparameter Optimization

    … under uncertainty is a fundamental problem encountered frequently in many real-world applications. This challenge has been rigorously formulated as the Stochastic Multi-Armed Bandit (SMAB) problem, which consists of a learner interacting with an environment. For each interaction, the …

    wustl Repository record for A Control Theoretic Approach to the Stochastic Multi-armed Bandit Problem With Applications in Hyperparameter Optimization (opens in a new tab)

  2. Private and Provably Efficient Federated Decision-Making

    In this thesis, we study sequential multi-armed bandit and reinforcement learning in the federated setting, where a group of agents collaborates to improve their collective reward by communicating over a network. We first study the multi-armed bandit problem in a decentralized environment. We study …

    mit Repository record for Private and Provably Efficient Federated Decision-Making (opens in a new tab)

  3. Robust sequential decision-making on networks

    In this thesis, I consider the research problem of designing optimal algorithms for two specific settings of the stochastic multi-armed bandit problem. The first setting considers the problem where rewards are drawn from a family of extremely heavy-tailed distributions known as a-stable …

    mit Repository record for Robust sequential decision-making on networks (opens in a new tab)

  4. Decomposition methods for large scale stochastic and robust optimization problems

    … families of stochastic and robust optimization problems in order to yield tractable approaches for large-scale real world application. We introduce a new type of a Markov decision problem named the Generalized Rest less Bandits Problem that encompasses a broad generalization of the restless …

    mit Repository record for Decomposition methods for large scale stochastic and robust optimization problems (opens in a new tab)

  5. LEARNING UNDER STRUCTURE AND UNCERTAINTY: ALGORITHMS FOR BANDIT AND ONLINE DECISION MAKING

    … with a particular focus on the multi-armed bandit and online learning frameworks. At the core of these settings lies a simple yet powerful interaction protocol: at each round the learner receives a question (information available prior to a decision), provides an answer (an action or …

    milano Repository record for LEARNING UNDER STRUCTURE AND UNCERTAINTY: ALGORITHMS FOR BANDIT AND ONLINE DECISION MAKING (opens in a new tab)

  6. Problem-Independent Regrets on Expectation-Dependent Multi-Armed Bandits

    … real world. We propose a new kind of multi-armed bandit problem where the expectation of outcomes may influence the agent’s utility which we call expectation-dependent multi-armed bandits and rationalize the choice of agents in Machina’s paradox lacking the IA. We design provably efficient …

    mit Repository record for Problem-Independent Regrets on Expectation-Dependent Multi-Armed Bandits (opens in a new tab)

  7. Analytics in promotional pricing and advertising

    … the right audience with their ads. We study this problem as a Multi-Armed Bandit problem with periodic budgets, and develop an Optimistic-Robust Learning algorithm with bounded expected regret. Practically, simulations on synthetic and real-world ad data show that the algorithm reduces regret by …

    mit Repository record for Analytics in promotional pricing and advertising (opens in a new tab)

  8. Low-complexity, low-regret link rate selection in rapidly varying wireless channels

    We consider the problem of transmitting at the optimal rate over a rapidly varying wireless channel with unknown statistics when the feedback about channel quality is very limited. One motivation for this problem is that, in emerging wireless networks, the use of mmWave bands means that the channel …

    uiuc Repository record for Low-complexity, low-regret link rate selection in rapidly varying wireless channels (opens in a new tab)

  9. Application of reinforcement learning methods to computer game dynamics

    … the simulation of several AS using 10-armed bandit problem averaged over 10000 epochs. The results show a considerable variation in performance in terms of latency and asymptotic direction. The Upper Confidence Bound comes out leader over most of the episode range, especially at about 100. …

    london-metro Repository record for Application of reinforcement learning methods to computer game dynamics (opens in a new tab)

  10. Particle Thompson sampling

    … Bayesian heuristic for solving stochastic bandit problems. But it is hard to implement in practice due to the intractability of maintaining a continuous posterior distribution. Particle Thompson sampling (PTS) is an approximation of Thompson sampling based on the simple idea of replacing …

    uiuc Repository record for Particle Thompson sampling (opens in a new tab)

  11. NON-STATIONARY MULTIARMED BANDITS FOR SATIATION AND SEASONALITY PHENOMENA IN MUSIC RECOMMENDER SYSTEMS

    … theoretical aspects of non-stationary multiarmed bandits, motivated by their application to music recommender systems. An intrinsic challenge of such systems lies in evolving user preferences. Rather than finding a single optimal item, the objective is to craft an ordered sequence of items …

    milano Repository record for NON-STATIONARY MULTIARMED BANDITS FOR SATIATION AND SEASONALITY PHENOMENA IN MUSIC RECOMMENDER SYSTEMS (opens in a new tab)

  12. ONLINE LEARNING, UNIFORM CONVERGENCE, AND A THEORY OF INTERPRETABILITY

    … of feedback models for multiple online learning problems, the sample complexity for uniform convergence, and a learning-theoretic approach to interpretable machine learning. First, we focus on online learning and investigate variants of the multi-armed bandit problem, including settings with …

    milano Repository record for ONLINE LEARNING, UNIFORM CONVERGENCE, AND A THEORY OF INTERPRETABILITY (opens in a new tab)

  13. ONLINE LEARNING WITH BANDITS FOR COVERAGE

    … of precision. In this work, we formulate this problem as the {\it online set coverage problem} and propose its solution for recommendation systems and the patrol assignment problem.</p> <p>We propose a novel online reinforcement learning algorithm inspired by the Multi-Armed Bandit problem to …

    syracuse-diss Repository record for ONLINE LEARNING WITH BANDITS FOR COVERAGE (opens in a new tab)

  14. Sequential decision making with feature-linear models

    This thesis is concerned with the problem of sequential decision making, where an agent interacts sequentially with an unknown environment and aims to maximise the sum of the rewards it receives. Our focus is on methods that model the reward as linear in some feature space. We consider a bandit

    cambridge Repository record for Sequential decision making with feature-linear models (opens in a new tab)

  15. Causal Structure Learning through Double Machine Learning

    … challenges and iv) the identifiability problem arises because multiple causal models can yield the same observational distribution, making it impossible to conclusively determine the true structure. In this thesis, we focus on the partial identification of underlying p causal structure …

    mit Repository record for Causal Structure Learning through Double Machine Learning (opens in a new tab)

  16. The exploration-exploitation trade-off in sequential decision making problems

    Sequential decision making problems require an agent to repeatedly choose between a series of actions. Common to such problems is the exploration-exploitation trade-off, where an agent must choose between the action expected to yield the best reward (exploitation) or trying an alternative action …

    lancaster Repository record for The exploration-exploitation trade-off in sequential decision making problems (opens in a new tab)

  17. Data-driven methods for personalized product recommendation systems

    … seller. Motivated by this increasingly prevalent problem, we propose an innovative model that selects, prices and recommends a personalized bundle of products to an online consumer. This model captures the trade-off between myopic profit maximization and inventory management, while selecting …

    mit Repository record for Data-driven methods for personalized product recommendation systems (opens in a new tab)

  18. Dynamic, data-driven decision-making in revenue management

    … Management (RM), this thesis studies various problems in sequential decision-making and demand learning. In the first module, we consider a personalized RM setting, where items with limited inventories are recommended to heterogeneous customers sequentially visiting an e-commerce platform. We …

    mit Repository record for Dynamic, data-driven decision-making in revenue management (opens in a new tab)

  19. Enhancing Capabilities of Assistive Robotic Arms: Learning, Control, and Object Manipulation

    … task learning as a sequence of multi-armed bandit problems, where each bandit problem corresponds to a waypoint in the robot's trajectory. We introduce an approximate posterior sampling solution that builds the robot's motion one waypoint at a time. Our simulations and real-world experiments …

    vt Repository record for Enhancing Capabilities of Assistive Robotic Arms: Learning, Control, and Object Manipulation (opens in a new tab)

  20. New Directions in Bandit Learning: Singularities and Random Walk Feedback

    <p>My thesis focuses new directions in bandit learning problems. In Chapter 1, I give an overview of the bandit learning literature, which lays the discussion framework for studies in Chapters 2 and 3. In Chapter 2, I study bandit learning problem in metric measure spaces. I start with multi-armed …

    duke Repository record for New Directions in Bandit Learning: Singularities and Random Walk Feedback (opens in a new tab)

Page 1 of 2