Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 58 for “"Multi-Armed Bandit"”.
-
A multi-armed bandit approach for batch mode active learning on information networks
… batch mode active learning algorithm, MABAL (Multi-Armed Bandit for Active Learning), for classification on heterogeneous information networks. Observing the parallels between active learning and multi-armed bandit (MAB), we base MABAL on an existing combinatorial MAB algorithm to combine …
-
Regulating exploration in multi-armed bandit problems with time patterns and dying arms
… before major holidays. The standard paradigm of multi-armed bandit analysis does not take these known patterns into account. This means that for applications in retail, where prices are fixed for periods of time, current bandit algorithms will not suffice. This work provides a framework and …
-
EFFECTS OF RESPONSE FREQUENCY CONSTRAINTS ON LEARNING IN A NON-STATIONARY MULTI-ARMED BANDIT TASK
An n-armed bandit task was used to investigate the trade-off between exploratory (choosing lesser-known options) and exploitive (choosing options with the greatest probability of reinforcement) human choice in a trial-and-error learning problem. In Experiment 1 a different probability of …
-
A Control Theoretic Approach to the Stochastic Multi-armed Bandit Problem With Applications in Hyperparameter Optimization
… has been rigorously formulated as the Stochastic Multi-Armed Bandit (SMAB) problem, which consists of a learner interacting with an environment. For each interaction, the learner selects an action and then receives a reward from the environment based on the chosen action. The learner's objective …
-
New Spatio-temporal Hawkes Process Models For Social Good
… that integrate Hawkes Process models with multi-armed bandit algorithms, high dimensional marks, and high-dimensional auxiliary data to solve problems in search and rescue, forecasting infectious disease, and early detection of overdose spikes. In Chapter 3, we develop a method applications …
-
Private and Provably Efficient Federated Decision-Making
In this thesis, we study sequential multi-armed bandit and reinforcement learning in the federated setting, where a group of agents collaborates to improve their collective reward by communicating over a network. We first study the multi-armed bandit problem in a decentralized environment. We study …
-
Autonomous adaptive acoustic relay positioning
… exploitation problem that is well-described by a multi-armed bandit formulation with an elegant solution in the form of Gittins indices. For an autonomous ocean vehicle traveling between distant waypoints, however, switching costs are significant. The multi-armed bandit with switching costs has no …
-
Multi-armed bandits and applications to large datasets
This thesis considers the multi-armed bandit (MAB) problem, both the traditional bandit feedback and graphical bandits when there is side information. Motivated by the Boltzmann exploration algorithm often used in the more general context of reinforcement learning, we present Almost Boltzmann …
-
A Bayesian bandit approach to personalized online coupon recommendations
… This paper resolves this problem by using a multi-armed bandit approach to balance the exploration (learning customers' preference for coupons) with exploitation (maximizing short term activation clicks). The proposed approach is evaluated with synthetic data. Results showed a 60% click lift …
-
Online advertisements and multi-armed bandits
We investigate a number of multi-armed bandit problems that model different aspects of online advertising, beginning with a survey of the key techniques that are commonly used to demonstrate the theoretical limitations and achievable results for the performance of multi-armed bandit algorithms. We …
-
Power Control and Resource Allocation for QoS-Constrained Wireless Networks
… such as machine-to-machine communications and multimedia services are placing growing demands on high-speed reliable transmissions and limited wireless spectrum resources. Although multiple-input multiple-output (MIMO) systems have shown the ability to provide reliable transmissions in fading …
-
BAYESIAN SEQUENTIAL OPTIMAL EXPERIMENTAL DESIGN FOR INVERSE PROBLEMS USING DEEP REINFORCEMENT LEARNING
… greedy design, black-box Bayesian optimization, multi-armed bandit optimization, dynamic programming, approximate dynamic programming, and reinforcement learning. This work showcases novel comparisons between the aforementioned methods and a new application of off-the-shelf reinforcement learning …
-
Problem-Independent Regrets on Expectation-Dependent Multi-Armed Bandits
… in the real world. We propose a new kind of multi-armed bandit problem where the expectation of outcomes may influence the agent’s utility which we call expectation-dependent multi-armed bandits and rationalize the choice of agents in Machina’s paradox lacking the IA. We design provably …
-
Analytics in promotional pricing and advertising
… with their ads. We study this problem as a Multi-Armed Bandit problem with periodic budgets, and develop an Optimistic-Robust Learning algorithm with bounded expected regret. Practically, simulations on synthetic and real-world ad data show that the algorithm reduces regret by at least …
-
Optimization for online platforms
… this problem, we consider a novel variant of the multi-armed bandit (MAB) problem, MAB with cost subsidy, which models many real-life applications where the learning agent has to pay to select an arm and is concerned about optimizing cumulative costs and rewards.
-
Exploration vs. exploitation in coupon personalization
… the retailers' offer allocation problem as a multi armed bandit and explore solution strategies.
-
Low-complexity, low-regret link rate selection in rapidly varying wireless channels
… Inspired by related problems in the context of multi-armed bandits, we consider a well-known algorithm called Thompson sampling to address this problem. However, unlike the traditional multi-armed bandit problem, a direct application of Thompson sampling results in a computational and storage …
-
Decentralized multi-user multi-armed bandits with user dependent reward distributions
… spectrum access problem is studied using a multi-player multi-armed bandits framework. We consider a decentralized multi-player stochastic multi-armed bandit model where the players cannot communicate with each other and can observe only their own actions and rewards. Furthermore, the …
-
Models of intelligence operations
… dynamic programming problem, namely the multi-armed bandit allocation (MABA) problem. The MABA framework models the efforts of a processor to search for intelligence items of the highest importance by making sequential samples from a collection of intelligence sources. Through Bayesian …
-
Machine learning blocks
… classification algorithms with a combination of Multi-Armed Bandit strategies and Gaussian Process optimizations, all in a distributed fashion in the cloud.
Page 1 of 3