Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 74 for “"bandits"”.

  1. Bandits in autoregressive Markov models

    Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-08-01

    uiuc Repository record for Bandits in autoregressive Markov models (opens in a new tab)

  2. ONLINE LEARNING WITH BANDITS FOR COVERAGE

    … can facilitate the coordination between bandits and therefore, reduce the overall complexity. Our graph-based bandit algorithm can select a much smaller set of items to cover a vast variety of users’ choices for recommendation systems. We present our experimental results in a partially …

    syracuse-diss Repository record for ONLINE LEARNING WITH BANDITS FOR COVERAGE (opens in a new tab)

  3. Online advertisements and multi-armed bandits

    … averages. Next, we consider multi-armed bandits with budgets, modeling how ad exchanges select which ad to display. We provide asymptotic regret lower bounds satisfied by any algorithm, and propose algorithms which match those lower bounds. We consider different types of budgets: …

    uiuc Repository record for Online advertisements and multi-armed bandits (opens in a new tab)

  4. Adversarial Bandits and which Leader to Follow

    L'abstract è presente nell'allegato / the abstract is in the attachment

    poli-torino Repository record for Adversarial Bandits and which Leader to Follow (opens in a new tab)

  5. Contextual Bandits with Neural Networks and Trees

    … με την εκμετάλλευση (exploitation). Τα bandits παρέχουν ένα απλό μοντέλο για αυτό το δίλημμα. Τα contextual bandits αποτελούν μια πολύ σημαντική κατηγορία, όπου ο πράκτορας(agent) έχει πρόσβαση σε πρόσθετες πληροφορίες που μπορεί να βοηθήσουν στην πρόβλεψη της ποιότητας των ενεργειών …

    athens Repository record for Contextual Bandits with Neural Networks and Trees (opens in a new tab)

  6. Optimizing deep learning networks using multi-armed bandits

    … those based on the following types of multi-arm bandits: (i) Epsilon-Greedy (ii) Upper Confidence Bounds (UCB) (iii) Thompson Sampling and (iv) Exponential Weight Algorithm for Exploration and Exploitation (EXP3). The algorithms were implemented in Python and a comprehensive empirical evaluation …

    salford Repository record for Optimizing deep learning networks using multi-armed bandits (opens in a new tab)

  7. New Models qnd Algorithms for Bandits and Markets

    Inspired by advertising markets, we consider large-scale sequential decision making problems in which a learner must deploy an algorithm to behave optimally under uncertainty. Although many of these problems can be modeled as contextual bandit problems, we argue that the tools and techniques for …

    penn Repository record for New Models qnd Algorithms for Bandits and Markets (opens in a new tab)

  8. Parametrized Stochastic Multi-armed Bandits with Binary Rewards

    … thesis, we consider the problem of multi-armed bandits with a large number of correlated arms. We assume that the arms have Bernoulli distributed rewards, independent across arms and across time, where the probabilities of success are parametrized by known attribute vectors for each arm, as well …

    uiuc Repository record for Parametrized Stochastic Multi-armed Bandits with Binary Rewards (opens in a new tab)

  9. Risk-averse multi-armed bandits and game theory

    … to study the fundamental limits of the existing bandits and game theory problems in a risk-averse framework and propose new ideas that address the shortcomings. The author believes that human beings are mostly risk-averse, so studying multi-armed bandits and game theory from the point of view of …

    uiuc Repository record for Risk-averse multi-armed bandits and game theory (opens in a new tab)

  10. Multi-armed bandits and applications to large datasets

    … the traditional bandit feedback and graphical bandits when there is side information. Motivated by the Boltzmann exploration algorithm often used in the more general context of reinforcement learning, we present Almost Boltzmann Exploration (ABE) which fixes the under-exploration issue while …

    uiuc Repository record for Multi-armed bandits and applications to large datasets (opens in a new tab)

  11. Problem-Independent Regrets on Expectation-Dependent Multi-Armed Bandits

    … which we call expectation-dependent multi-armed bandits and rationalize the choice of agents in Machina’s paradox lacking the IA. We design provably efficient algorithms with low minimax regrets and show their consistency of time horizon T with corresponding regret lower bounds, revealing …

    mit Repository record for Problem-Independent Regrets on Expectation-Dependent Multi-Armed Bandits (opens in a new tab)

  12. Topics in high-dimensional linear bandits and approximate Bayesian sampling

    Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2024-05-01

    uiuc Repository record for Topics in high-dimensional linear bandits and approximate Bayesian sampling (opens in a new tab)

  13. Learning Heterogeneous Resource-Constrained Task Allocation Using Concurrent Multi-Task Bandits

    … algorithm called Concurrent Multi-Task Adaptive Bandits (CMTAB), which leverages and builds upon continuum-armed bandit algorithms. Our experiments, which involve detailed numerical simulations and a simulated emergency response task, demonstrate that CMTAB is effective at balancing exploration …

    gatech Repository record for Learning Heterogeneous Resource-Constrained Task Allocation Using Concurrent Multi-Task Bandits (opens in a new tab)

  14. “Suppressing bandits”: Social policing of warlord government in Manchuria, 1918-1928

    Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms

    uiuc Repository record for “Suppressing bandits”: Social policing of warlord government in Manchuria, 1918-1928 (opens in a new tab)

  15. Decentralized multi-user multi-armed bandits with user dependent reward distributions

    … is studied using a multi-player multi-armed bandits framework. We consider a decentralized multi-player stochastic multi-armed bandit model where the players cannot communicate with each other and can observe only their own actions and rewards. Furthermore, the environment may appear …

    uiuc Repository record for Decentralized multi-user multi-armed bandits with user dependent reward distributions (opens in a new tab)

  16. Dealers, insiders and bandits : learning and its effects on market outcomes

    This thesis seeks to contribute to the understanding of markets populated by boundedly rational agents who learn from experience. Bounded rationality and learning have both been the focus of much research in computer science, economics and finance theory. However, we are at a critical stage in …

    mit Repository record for Dealers, insiders and bandits : learning and its effects on market outcomes (opens in a new tab)

  17. Bayesian tuning and bandits : an extensible, open source library for AutoML

    The goal of this thesis is to build an extensible and open source library that handles the problems of tuning the hyperparameters of a machine learning pipeline, selecting between multiple pipelines, and recommending a pipeline. We devise a library that users can integrate into their existing …

    mit Repository record for Bayesian tuning and bandits : an extensible, open source library for AutoML (opens in a new tab)

  18. NON-STATIONARY MULTIARMED BANDITS FOR SATIATION AND SEASONALITY PHENOMENA IN MUSIC RECOMMENDER SYSTEMS

    … theoretical aspects of non-stationary multiarmed bandits, motivated by their application to music recommender systems. An intrinsic challenge of such systems lies in evolving user preferences. Rather than finding a single optimal item, the objective is to craft an ordered sequence of items …

    milano Repository record for NON-STATIONARY MULTIARMED BANDITS FOR SATIATION AND SEASONALITY PHENOMENA IN MUSIC RECOMMENDER SYSTEMS (opens in a new tab)

  19. Adaptive Preference Learning With Bandit Feedback: Information Filtering, Dueling Bandits and Incentivizing Exploration

    … to existing work on classical multi-armed bandits, dueling bandits, and incentivizing exploration. For each type of feedback and application setting, we provide an algorithm and a theoretical analysis bounding its regret. We demonstrate through numerical experiments that our algorithms …

    cornell Repository record for Adaptive Preference Learning With Bandit Feedback: Information Filtering, Dueling Bandits and Incentivizing Exploration (opens in a new tab)

  20. Theory and Practice of Large-scale Logistics: Offline Contextual Bandits and Decomposition Methods

    … The first part focuses on offline contextual bandits, drawing from work at Uber on driver incentive programs. I introduce Empirical Soft Regret (ESR), a novel loss function for value-based learning that addresses limitations of accuracy-based approaches in misspecified settings. Unlike …

    cornell Repository record for Theory and Practice of Large-scale Logistics: Offline Contextual Bandits and Decomposition Methods (opens in a new tab)

Page 1 of 4