Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 29 for “"multi-armed bandits"”.
-
Online advertisements and multi-armed bandits
We investigate a number of multi-armed bandit problems that model different aspects of online advertising, beginning with a survey of the key techniques that are commonly used to demonstrate the theoretical limitations and achievable results for the performance of multi-armed bandit algorithms. We …
-
Optimizing deep learning networks using multi-armed bandits
… for pruning that utilize a framework, known as a multi-armed bandit, which has been successfully applied in applications where there is a need to learn which option to select given the outcome of trials. There are several different multi-arm bandit methods, and these have been used to develop new …
-
Parametrized Stochastic Multi-armed Bandits with Binary Rewards
In this thesis, we consider the problem of multi-armed bandits with a large number of correlated arms. We assume that the arms have Bernoulli distributed rewards, independent across arms and across time, where the probabilities of success are parametrized by known attribute vectors for each arm, as …
-
Risk-averse multi-armed bandits and game theory
The multi-armed bandit (MAB) and game theory literature is mainly focused on the expected cumulative reward and the expected payoffs in a game, respectively. In contrast, the rewards and the payoffs are often random variables whose expected values only capture a vague idea of the overall …
-
Multi-armed bandits and applications to large datasets
This thesis considers the multi-armed bandit (MAB) problem, both the traditional bandit feedback and graphical bandits when there is side information. Motivated by the Boltzmann exploration algorithm often used in the more general context of reinforcement learning, we present Almost Boltzmann …
-
Problem-Independent Regrets on Expectation-Dependent Multi-Armed Bandits
… in the real world. We propose a new kind of multi-armed bandit problem where the expectation of outcomes may influence the agent’s utility which we call expectation-dependent multi-armed bandits and rationalize the choice of agents in Machina’s paradox lacking the IA. We design provably …
-
Decentralized multi-user multi-armed bandits with user dependent reward distributions
… spectrum access problem is studied using a multi-player multi-armed bandits framework. We consider a decentralized multi-player stochastic multi-armed bandit model where the players cannot communicate with each other and can observe only their own actions and rewards. Furthermore, the …
-
On upper confidence bound algorithms for piecewise-stationary stochastic multi-armed bandits and the variants
In recent years, multi-armed bandit (MAB) problems have received much attention, as they model many real-world applications such as online recommendation, web search and crowdsourcing tasks. The core of MAB algorithms is addressing the exploration versus exploitation dilemma, and finding the right …
-
Adaptive Preference Learning With Bandit Feedback: Information Filtering, Dueling Bandits and Incentivizing Exploration
… respectively to existing work on classical multi-armed bandits, dueling bandits, and incentivizing exploration. For each type of feedback and application setting, we provide an algorithm and a theoretical analysis bounding its regret. We demonstrate through numerical experiments that our …
-
An Introduction to Reinforcement Learning
… in basic probability, linear algebra, and multivariable calculus. We systematically cover the fundamentals of Markov decision processes, optimal control, multi-armed bandits, fitted dynamic programming algorithms, policy gradient methods, imitation learning, and tree search-based planning …
-
Low-complexity, low-regret link rate selection in rapidly varying wireless channels
… Inspired by related problems in the context of multi-armed bandits, we consider a well-known algorithm called Thompson sampling to address this problem. However, unlike the traditional multi-armed bandit problem, a direct application of Thompson sampling results in a computational and storage …
-
Decision-Making Under Uncertainty: From Theory to Practice
… framework with a case study on contextual bandits for warfarin dosing where we are concerned with the cost of exploration across multiple races and age groups. Next, we study the classical problem of minimizing regret for multi-armed bandits. In this classic problem, there are several …
-
Sample-efficient reinforcement learning
… The first problem we consider is the structured multi-armed bandits problem, motivated by an application in wireless networks. The second problem we consider is the bandits with two-level feedback problem, motivated by an application in panoramic video streaming. The third problem we consider is …
-
The Modeling Spectrum of Data-Driven Decision Making
… good farming practices in the framework of multi-armed bandits with expert advice. We extend the setting from finitely many experts to any countably infinite set and provide algorithms that are provably optimal. Second, we explore optimizing perturbations for cell reprogramming in batched …
-
Democratizing data science through interactive curation of ML pipelines
… and pruning strategies combining cost-based Multi-Armed Bandits and Bayesian Optimization. We evaluate the system on over 300 datasets and compare against other AutoML tools, including the current NIPS winner, as well as expert solutions. Not only is Alpine Meadow able to significantly …
-
Certifying robustness in inference and learning problems
… detection robust to distribution shifts, and multi-player multi-armed bandits robust to adversarial attacks. Principled approaches for these problems with theoretical guarantees are derived using tools from statistics, information theory and optimization, that are practical, resilient and …
-
DISTRIBUTED AND DELAYED ONLINE LEARNING
… in settings with partial feedback (e.g., multi-armed bandits) by developing a meta-algorithm that achieves near-optimal regret with significantly reduced sensitivity to total delay. Second, we consider distributed online convex optimization over communication graphs, in which a network of …
-
On the Sample Efficiency of Data-Driven Decision Making
… interactive decision making as exemplified by multi-armed bandits and reinforcement learning. The first part of the thesis develops novel algorithmic and theoretical tools to enhance decision making in these regimes and to bridge the gaps between them. We revisit logistic regression in the …
-
Reinforcement Learning-based Optimization of Multiple Access in Wireless Networks
In this thesis, we study the problem of Multiple Access (MA) in wireless networks and design adaptive solutions based on Reinforcement Learning (RL). We analyze the importance of MA in the current communications scenery, where bandwidth-hungry applications emerge due to the co-evolution of …
-
Predictive and Prescriptive Trees for Optimization and Control Problems
… learning approach to the optimal control of multiclass fluid queueing networks (MFQNETs). We prove that a piecewise constant optimal policy exists for MFQNET control problems, with segments separated by hyperplanes passing through the origin. We use Optimal Classification Trees with …
Page 1 of 2