Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 29 for “"multi-armed bandits"”.

  1. Online advertisements and multi-armed bandits

    We investigate a number of multi-armed bandit problems that model different aspects of online advertising, beginning with a survey of the key techniques that are commonly used to demonstrate the theoretical limitations and achievable results for the performance of multi-armed bandit algorithms. We …

    uiuc Repository record for Online advertisements and multi-armed bandits (opens in a new tab)

  2. Optimizing deep learning networks using multi-armed bandits

    … for pruning that utilize a framework, known as a multi-armed bandit, which has been successfully applied in applications where there is a need to learn which option to select given the outcome of trials. There are several different multi-arm bandit methods, and these have been used to develop new …

    salford Repository record for Optimizing deep learning networks using multi-armed bandits (opens in a new tab)

  3. Parametrized Stochastic Multi-armed Bandits with Binary Rewards

    In this thesis, we consider the problem of multi-armed bandits with a large number of correlated arms. We assume that the arms have Bernoulli distributed rewards, independent across arms and across time, where the probabilities of success are parametrized by known attribute vectors for each arm, as …

    uiuc Repository record for Parametrized Stochastic Multi-armed Bandits with Binary Rewards (opens in a new tab)

  4. Risk-averse multi-armed bandits and game theory

    The multi-armed bandit (MAB) and game theory literature is mainly focused on the expected cumulative reward and the expected payoffs in a game, respectively. In contrast, the rewards and the payoffs are often random variables whose expected values only capture a vague idea of the overall …

    uiuc Repository record for Risk-averse multi-armed bandits and game theory (opens in a new tab)

  5. Multi-armed bandits and applications to large datasets

    This thesis considers the multi-armed bandit (MAB) problem, both the traditional bandit feedback and graphical bandits when there is side information. Motivated by the Boltzmann exploration algorithm often used in the more general context of reinforcement learning, we present Almost Boltzmann …

    uiuc Repository record for Multi-armed bandits and applications to large datasets (opens in a new tab)

  6. Problem-Independent Regrets on Expectation-Dependent Multi-Armed Bandits

    … in the real world. We propose a new kind of multi-armed bandit problem where the expectation of outcomes may influence the agent’s utility which we call expectation-dependent multi-armed bandits and rationalize the choice of agents in Machina’s paradox lacking the IA. We design provably …

    mit Repository record for Problem-Independent Regrets on Expectation-Dependent Multi-Armed Bandits (opens in a new tab)

  7. Decentralized multi-user multi-armed bandits with user dependent reward distributions

    … spectrum access problem is studied using a multi-player multi-armed bandits framework. We consider a decentralized multi-player stochastic multi-armed bandit model where the players cannot communicate with each other and can observe only their own actions and rewards. Furthermore, the …

    uiuc Repository record for Decentralized multi-user multi-armed bandits with user dependent reward distributions (opens in a new tab)

  8. On upper confidence bound algorithms for piecewise-stationary stochastic multi-armed bandits and the variants

    In recent years, multi-armed bandit (MAB) problems have received much attention, as they model many real-world applications such as online recommendation, web search and crowdsourcing tasks. The core of MAB algorithms is addressing the exploration versus exploitation dilemma, and finding the right …

    uiuc Repository record for On upper confidence bound algorithms for piecewise-stationary stochastic multi-armed bandits and the variants (opens in a new tab)

  9. Adaptive Preference Learning With Bandit Feedback: Information Filtering, Dueling Bandits and Incentivizing Exploration

    … respectively to existing work on classical multi-armed bandits, dueling bandits, and incentivizing exploration. For each type of feedback and application setting, we provide an algorithm and a theoretical analysis bounding its regret. We demonstrate through numerical experiments that our …

    cornell Repository record for Adaptive Preference Learning With Bandit Feedback: Information Filtering, Dueling Bandits and Incentivizing Exploration (opens in a new tab)

  10. An Introduction to Reinforcement Learning

    … in basic probability, linear algebra, and multivariable calculus. We systematically cover the fundamentals of Markov decision processes, optimal control, multi-armed bandits, fitted dynamic programming algorithms, policy gradient methods, imitation learning, and tree search-based planning …

    harvard Repository record for An Introduction to Reinforcement Learning (opens in a new tab)

  11. Low-complexity, low-regret link rate selection in rapidly varying wireless channels

    … Inspired by related problems in the context of multi-armed bandits, we consider a well-known algorithm called Thompson sampling to address this problem. However, unlike the traditional multi-armed bandit problem, a direct application of Thompson sampling results in a computational and storage …

    uiuc Repository record for Low-complexity, low-regret link rate selection in rapidly varying wireless channels (opens in a new tab)

  12. Decision-Making Under Uncertainty: From Theory to Practice

    … framework with a case study on contextual bandits for warfarin dosing where we are concerned with the cost of exploration across multiple races and age groups. Next, we study the classical problem of minimizing regret for multi-armed bandits. In this classic problem, there are several …

    mit Repository record for Decision-Making Under Uncertainty: From Theory to Practice (opens in a new tab)

  13. Sample-efficient reinforcement learning

    … The first problem we consider is the structured multi-armed bandits problem, motivated by an application in wireless networks. The second problem we consider is the bandits with two-level feedback problem, motivated by an application in panoramic video streaming. The third problem we consider is …

    uiuc Repository record for Sample-efficient reinforcement learning (opens in a new tab)

  14. The Modeling Spectrum of Data-Driven Decision Making

    … good farming practices in the framework of multi-armed bandits with expert advice. We extend the setting from finitely many experts to any countably infinite set and provide algorithms that are provably optimal. Second, we explore optimizing perturbations for cell reprogramming in batched …

    mit Repository record for The Modeling Spectrum of Data-Driven Decision Making (opens in a new tab)

  15. Democratizing data science through interactive curation of ML pipelines

    … and pruning strategies combining cost-based Multi-Armed Bandits and Bayesian Optimization. We evaluate the system on over 300 datasets and compare against other AutoML tools, including the current NIPS winner, as well as expert solutions. Not only is Alpine Meadow able to significantly …

    mit Repository record for Democratizing data science through interactive curation of ML pipelines (opens in a new tab)

  16. Certifying robustness in inference and learning problems

    … detection robust to distribution shifts, and multi-player multi-armed bandits robust to adversarial attacks. Principled approaches for these problems with theoretical guarantees are derived using tools from statistics, information theory and optimization, that are practical, resilient and …

    uiuc Repository record for Certifying robustness in inference and learning problems (opens in a new tab)

  17. DISTRIBUTED AND DELAYED ONLINE LEARNING

    … in settings with partial feedback (e.g., multi-armed bandits) by developing a meta-algorithm that achieves near-optimal regret with significantly reduced sensitivity to total delay. Second, we consider distributed online convex optimization over communication graphs, in which a network of …

    milano Repository record for DISTRIBUTED AND DELAYED ONLINE LEARNING (opens in a new tab)

  18. On the Sample Efficiency of Data-Driven Decision Making

    … interactive decision making as exemplified by multi-armed bandits and reinforcement learning. The first part of the thesis develops novel algorithmic and theoretical tools to enhance decision making in these regimes and to bridge the gaps between them. We revisit logistic regression in the …

    mit Repository record for On the Sample Efficiency of Data-Driven Decision Making (opens in a new tab)

  19. Reinforcement Learning-based Optimization of Multiple Access in Wireless Networks

    In this thesis, we study the problem of Multiple Access (MA) in wireless networks and design adaptive solutions based on Reinforcement Learning (RL). We analyze the importance of MA in the current communications scenery, where bandwidth-hungry applications emerge due to the co-evolution of …

    essex Repository record for Reinforcement Learning-based Optimization of Multiple Access in Wireless Networks (opens in a new tab)

  20. Predictive and Prescriptive Trees for Optimization and Control Problems

    … learning approach to the optimal control of multiclass fluid queueing networks (MFQNETs). We prove that a piecewise constant optimal policy exists for MFQNET control problems, with segments separated by hyperplanes passing through the origin. We use Optimal Classification Trees with …

    mit Repository record for Predictive and Prescriptive Trees for Optimization and Control Problems (opens in a new tab)

Page 1 of 2