Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 25 for “"bandit problem"”.
-
A Control Theoretic Approach to the Stochastic Multi-armed Bandit Problem With Applications in Hyperparameter Optimization
… under uncertainty is a fundamental problem encountered frequently in many real-world applications. This challenge has been rigorously formulated as the Stochastic Multi-Armed Bandit (SMAB) problem, which consists of a learner interacting with an environment. For each interaction, the …
-
Private and Provably Efficient Federated Decision-Making
In this thesis, we study sequential multi-armed bandit and reinforcement learning in the federated setting, where a group of agents collaborates to improve their collective reward by communicating over a network. We first study the multi-armed bandit problem in a decentralized environment. We study …
-
Robust sequential decision-making on networks
In this thesis, I consider the research problem of designing optimal algorithms for two specific settings of the stochastic multi-armed bandit problem. The first setting considers the problem where rewards are drawn from a family of extremely heavy-tailed distributions known as a-stable …
-
Decomposition methods for large scale stochastic and robust optimization problems
… families of stochastic and robust optimization problems in order to yield tractable approaches for large-scale real world application. We introduce a new type of a Markov decision problem named the Generalized Rest less Bandits Problem that encompasses a broad generalization of the restless …
-
LEARNING UNDER STRUCTURE AND UNCERTAINTY: ALGORITHMS FOR BANDIT AND ONLINE DECISION MAKING
… with a particular focus on the multi-armed bandit and online learning frameworks. At the core of these settings lies a simple yet powerful interaction protocol: at each round the learner receives a question (information available prior to a decision), provides an answer (an action or …
-
Problem-Independent Regrets on Expectation-Dependent Multi-Armed Bandits
… real world. We propose a new kind of multi-armed bandit problem where the expectation of outcomes may influence the agent’s utility which we call expectation-dependent multi-armed bandits and rationalize the choice of agents in Machina’s paradox lacking the IA. We design provably efficient …
-
Analytics in promotional pricing and advertising
… the right audience with their ads. We study this problem as a Multi-Armed Bandit problem with periodic budgets, and develop an Optimistic-Robust Learning algorithm with bounded expected regret. Practically, simulations on synthetic and real-world ad data show that the algorithm reduces regret by …
-
Low-complexity, low-regret link rate selection in rapidly varying wireless channels
We consider the problem of transmitting at the optimal rate over a rapidly varying wireless channel with unknown statistics when the feedback about channel quality is very limited. One motivation for this problem is that, in emerging wireless networks, the use of mmWave bands means that the channel …
-
Application of reinforcement learning methods to computer game dynamics
… the simulation of several AS using 10-armed bandit problem averaged over 10000 epochs. The results show a considerable variation in performance in terms of latency and asymptotic direction. The Upper Confidence Bound comes out leader over most of the episode range, especially at about 100. …
-
Particle Thompson sampling
… Bayesian heuristic for solving stochastic bandit problems. But it is hard to implement in practice due to the intractability of maintaining a continuous posterior distribution. Particle Thompson sampling (PTS) is an approximation of Thompson sampling based on the simple idea of replacing …
-
NON-STATIONARY MULTIARMED BANDITS FOR SATIATION AND SEASONALITY PHENOMENA IN MUSIC RECOMMENDER SYSTEMS
… theoretical aspects of non-stationary multiarmed bandits, motivated by their application to music recommender systems. An intrinsic challenge of such systems lies in evolving user preferences. Rather than finding a single optimal item, the objective is to craft an ordered sequence of items …
-
ONLINE LEARNING, UNIFORM CONVERGENCE, AND A THEORY OF INTERPRETABILITY
… of feedback models for multiple online learning problems, the sample complexity for uniform convergence, and a learning-theoretic approach to interpretable machine learning. First, we focus on online learning and investigate variants of the multi-armed bandit problem, including settings with …
-
ONLINE LEARNING WITH BANDITS FOR COVERAGE
… of precision. In this work, we formulate this problem as the {\it online set coverage problem} and propose its solution for recommendation systems and the patrol assignment problem.</p> <p>We propose a novel online reinforcement learning algorithm inspired by the Multi-Armed Bandit problem to …
-
Sequential decision making with feature-linear models
This thesis is concerned with the problem of sequential decision making, where an agent interacts sequentially with an unknown environment and aims to maximise the sum of the rewards it receives. Our focus is on methods that model the reward as linear in some feature space. We consider a bandit …
-
Causal Structure Learning through Double Machine Learning
… challenges and iv) the identifiability problem arises because multiple causal models can yield the same observational distribution, making it impossible to conclusively determine the true structure. In this thesis, we focus on the partial identification of underlying p causal structure …
-
The exploration-exploitation trade-off in sequential decision making problems
Sequential decision making problems require an agent to repeatedly choose between a series of actions. Common to such problems is the exploration-exploitation trade-off, where an agent must choose between the action expected to yield the best reward (exploitation) or trying an alternative action …
-
Data-driven methods for personalized product recommendation systems
… seller. Motivated by this increasingly prevalent problem, we propose an innovative model that selects, prices and recommends a personalized bundle of products to an online consumer. This model captures the trade-off between myopic profit maximization and inventory management, while selecting …
-
Dynamic, data-driven decision-making in revenue management
… Management (RM), this thesis studies various problems in sequential decision-making and demand learning. In the first module, we consider a personalized RM setting, where items with limited inventories are recommended to heterogeneous customers sequentially visiting an e-commerce platform. We …
-
Enhancing Capabilities of Assistive Robotic Arms: Learning, Control, and Object Manipulation
… task learning as a sequence of multi-armed bandit problems, where each bandit problem corresponds to a waypoint in the robot's trajectory. We introduce an approximate posterior sampling solution that builds the robot's motion one waypoint at a time. Our simulations and real-world experiments …
-
New Directions in Bandit Learning: Singularities and Random Walk Feedback
<p>My thesis focuses new directions in bandit learning problems. In Chapter 1, I give an overview of the bandit learning literature, which lays the discussion framework for studies in Chapters 2 and 3. In Chapter 2, I study bandit learning problem in metric measure spaces. I start with multi-armed …
Page 1 of 2