Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 58 for “"Multi Armed Bandit"”.

  1. A multi-armed bandit approach for batch mode active learning on information networks

    … batch mode active learning algorithm, MABAL (Multi-Armed Bandit for Active Learning), for classification on heterogeneous information networks. Observing the parallels between active learning and multi-armed bandit (MAB), we base MABAL on an existing combinatorial MAB algorithm to combine …

    uiuc Repository record for A multi-armed bandit approach for batch mode active learning on information networks (opens in a new tab)

  2. Regulating exploration in multi-armed bandit problems with time patterns and dying arms

    … before major holidays. The standard paradigm of multi-armed bandit analysis does not take these known patterns into account. This means that for applications in retail, where prices are fixed for periods of time, current bandit algorithms will not suffice. This work provides a framework and …

    mit Repository record for Regulating exploration in multi-armed bandit problems with time patterns and dying arms (opens in a new tab)

  3. EFFECTS OF RESPONSE FREQUENCY CONSTRAINTS ON LEARNING IN A NON-STATIONARY MULTI-ARMED BANDIT TASK

    An n-armed bandit task was used to investigate the trade-off between exploratory (choosing lesser-known options) and exploitive (choosing options with the greatest probability of reinforcement) human choice in a trial-and-error learning problem. In Experiment 1 a different probability of …

    siu-theses Repository record for EFFECTS OF RESPONSE FREQUENCY CONSTRAINTS ON LEARNING IN A NON-STATIONARY MULTI-ARMED BANDIT TASK (opens in a new tab)

  4. A Control Theoretic Approach to the Stochastic Multi-armed Bandit Problem With Applications in Hyperparameter Optimization

    … has been rigorously formulated as the Stochastic Multi-Armed Bandit (SMAB) problem, which consists of a learner interacting with an environment. For each interaction, the learner selects an action and then receives a reward from the environment based on the chosen action. The learner's objective …

    wustl Repository record for A Control Theoretic Approach to the Stochastic Multi-armed Bandit Problem With Applications in Hyperparameter Optimization (opens in a new tab)

  5. New Spatio-temporal Hawkes Process Models For Social Good

    … that integrate Hawkes Process models with multi-armed bandit algorithms, high dimensional marks, and high-dimensional auxiliary data to solve problems in search and rescue, forecasting infectious disease, and early detection of overdose spikes. In Chapter 3, we develop a method applications …

    iupui Repository record for New Spatio-temporal Hawkes Process Models For Social Good (opens in a new tab)

  6. Private and Provably Efficient Federated Decision-Making

    In this thesis, we study sequential multi-armed bandit and reinforcement learning in the federated setting, where a group of agents collaborates to improve their collective reward by communicating over a network. We first study the multi-armed bandit problem in a decentralized environment. We study …

    mit Repository record for Private and Provably Efficient Federated Decision-Making (opens in a new tab)

  7. Autonomous adaptive acoustic relay positioning

    … exploitation problem that is well-described by a multi-armed bandit formulation with an elegant solution in the form of Gittins indices. For an autonomous ocean vehicle traveling between distant waypoints, however, switching costs are significant. The multi-armed bandit with switching costs has no …

    mit Repository record for Autonomous adaptive acoustic relay positioning (opens in a new tab)

  8. Multi-armed bandits and applications to large datasets

    This thesis considers the multi-armed bandit (MAB) problem, both the traditional bandit feedback and graphical bandits when there is side information. Motivated by the Boltzmann exploration algorithm often used in the more general context of reinforcement learning, we present Almost Boltzmann …

    uiuc Repository record for Multi-armed bandits and applications to large datasets (opens in a new tab)

  9. A Bayesian bandit approach to personalized online coupon recommendations

    … This paper resolves this problem by using a multi-armed bandit approach to balance the exploration (learning customers' preference for coupons) with exploitation (maximizing short term activation clicks). The proposed approach is evaluated with synthetic data. Results showed a 60% click lift …

    mit Repository record for A Bayesian bandit approach to personalized online coupon recommendations (opens in a new tab)

  10. Online advertisements and multi-armed bandits

    We investigate a number of multi-armed bandit problems that model different aspects of online advertising, beginning with a survey of the key techniques that are commonly used to demonstrate the theoretical limitations and achievable results for the performance of multi-armed bandit algorithms. We …

    uiuc Repository record for Online advertisements and multi-armed bandits (opens in a new tab)

  11. Power Control and Resource Allocation for QoS-Constrained Wireless Networks

    … such as machine-to-machine communications and multimedia services are placing growing demands on high-speed reliable transmissions and limited wireless spectrum resources. Although multiple-input multiple-output (MIMO) systems have shown the ability to provide reliable transmissions in fading …

    cambridge Repository record for Power Control and Resource Allocation for QoS-Constrained Wireless Networks (opens in a new tab)

  12. BAYESIAN SEQUENTIAL OPTIMAL EXPERIMENTAL DESIGN FOR INVERSE PROBLEMS USING DEEP REINFORCEMENT LEARNING

    … greedy design, black-box Bayesian optimization, multi-armed bandit optimization, dynamic programming, approximate dynamic programming, and reinforcement learning. This work showcases novel comparisons between the aforementioned methods and a new application of off-the-shelf reinforcement learning …

    umn Repository record for BAYESIAN SEQUENTIAL OPTIMAL EXPERIMENTAL DESIGN FOR INVERSE PROBLEMS USING DEEP REINFORCEMENT LEARNING (opens in a new tab)

  13. Problem-Independent Regrets on Expectation-Dependent Multi-Armed Bandits

    … in the real world. We propose a new kind of multi-armed bandit problem where the expectation of outcomes may influence the agent’s utility which we call expectation-dependent multi-armed bandits and rationalize the choice of agents in Machina’s paradox lacking the IA. We design provably …

    mit Repository record for Problem-Independent Regrets on Expectation-Dependent Multi-Armed Bandits (opens in a new tab)

  14. Analytics in promotional pricing and advertising

    … with their ads. We study this problem as a Multi-Armed Bandit problem with periodic budgets, and develop an Optimistic-Robust Learning algorithm with bounded expected regret. Practically, simulations on synthetic and real-world ad data show that the algorithm reduces regret by at least …

    mit Repository record for Analytics in promotional pricing and advertising (opens in a new tab)

  15. Optimization for online platforms

    … this problem, we consider a novel variant of the multi-armed bandit (MAB) problem, MAB with cost subsidy, which models many real-life applications where the learning agent has to pay to select an arm and is concerned about optimizing cumulative costs and rewards.

    mit Repository record for Optimization for online platforms (opens in a new tab)

  16. Exploration vs. exploitation in coupon personalization

    … the retailers' offer allocation problem as a multi armed bandit and explore solution strategies.

    mit Repository record for Exploration vs. exploitation in coupon personalization (opens in a new tab)

  17. Low-complexity, low-regret link rate selection in rapidly varying wireless channels

    … Inspired by related problems in the context of multi-armed bandits, we consider a well-known algorithm called Thompson sampling to address this problem. However, unlike the traditional multi-armed bandit problem, a direct application of Thompson sampling results in a computational and storage …

    uiuc Repository record for Low-complexity, low-regret link rate selection in rapidly varying wireless channels (opens in a new tab)

  18. Decentralized multi-user multi-armed bandits with user dependent reward distributions

    … spectrum access problem is studied using a multi-player multi-armed bandits framework. We consider a decentralized multi-player stochastic multi-armed bandit model where the players cannot communicate with each other and can observe only their own actions and rewards. Furthermore, the …

    uiuc Repository record for Decentralized multi-user multi-armed bandits with user dependent reward distributions (opens in a new tab)

  19. Models of intelligence operations

    … dynamic programming problem, namely the multi-armed bandit allocation (MABA) problem. The MABA framework models the efforts of a processor to search for intelligence items of the highest importance by making sequential samples from a collection of intelligence sources. Through Bayesian …

    lancaster Repository record for Models of intelligence operations (opens in a new tab)

  20. Machine learning blocks

    … classification algorithms with a combination of Multi-Armed Bandit strategies and Gaussian Process optimizations, all in a distributed fashion in the cloud.

    mit Repository record for Machine learning blocks (opens in a new tab)

Page 1 of 3