Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 74 for “"bandits"”.
-
Bandits in autoregressive Markov models
Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-08-01
-
ONLINE LEARNING WITH BANDITS FOR COVERAGE
… can facilitate the coordination between bandits and therefore, reduce the overall complexity. Our graph-based bandit algorithm can select a much smaller set of items to cover a vast variety of users’ choices for recommendation systems. We present our experimental results in a partially …
-
Online advertisements and multi-armed bandits
… averages. Next, we consider multi-armed bandits with budgets, modeling how ad exchanges select which ad to display. We provide asymptotic regret lower bounds satisfied by any algorithm, and propose algorithms which match those lower bounds. We consider different types of budgets: …
-
Adversarial Bandits and which Leader to Follow
L'abstract è presente nell'allegato / the abstract is in the attachment
-
Contextual Bandits with Neural Networks and Trees
… με την εκμετάλλευση (exploitation). Τα bandits παρέχουν ένα απλό μοντέλο για αυτό το δίλημμα. Τα contextual bandits αποτελούν μια πολύ σημαντική κατηγορία, όπου ο πράκτορας(agent) έχει πρόσβαση σε πρόσθετες πληροφορίες που μπορεί να βοηθήσουν στην πρόβλεψη της ποιότητας των ενεργειών …
-
Optimizing deep learning networks using multi-armed bandits
… those based on the following types of multi-arm bandits: (i) Epsilon-Greedy (ii) Upper Confidence Bounds (UCB) (iii) Thompson Sampling and (iv) Exponential Weight Algorithm for Exploration and Exploitation (EXP3). The algorithms were implemented in Python and a comprehensive empirical evaluation …
-
New Models qnd Algorithms for Bandits and Markets
Inspired by advertising markets, we consider large-scale sequential decision making problems in which a learner must deploy an algorithm to behave optimally under uncertainty. Although many of these problems can be modeled as contextual bandit problems, we argue that the tools and techniques for …
-
Parametrized Stochastic Multi-armed Bandits with Binary Rewards
… thesis, we consider the problem of multi-armed bandits with a large number of correlated arms. We assume that the arms have Bernoulli distributed rewards, independent across arms and across time, where the probabilities of success are parametrized by known attribute vectors for each arm, as well …
-
Risk-averse multi-armed bandits and game theory
… to study the fundamental limits of the existing bandits and game theory problems in a risk-averse framework and propose new ideas that address the shortcomings. The author believes that human beings are mostly risk-averse, so studying multi-armed bandits and game theory from the point of view of …
-
Multi-armed bandits and applications to large datasets
… the traditional bandit feedback and graphical bandits when there is side information. Motivated by the Boltzmann exploration algorithm often used in the more general context of reinforcement learning, we present Almost Boltzmann Exploration (ABE) which fixes the under-exploration issue while …
-
Problem-Independent Regrets on Expectation-Dependent Multi-Armed Bandits
… which we call expectation-dependent multi-armed bandits and rationalize the choice of agents in Machina’s paradox lacking the IA. We design provably efficient algorithms with low minimax regrets and show their consistency of time horizon T with corresponding regret lower bounds, revealing …
-
Topics in high-dimensional linear bandits and approximate Bayesian sampling
Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2024-05-01
-
Learning Heterogeneous Resource-Constrained Task Allocation Using Concurrent Multi-Task Bandits
… algorithm called Concurrent Multi-Task Adaptive Bandits (CMTAB), which leverages and builds upon continuum-armed bandit algorithms. Our experiments, which involve detailed numerical simulations and a simulated emergency response task, demonstrate that CMTAB is effective at balancing exploration …
-
“Suppressing bandits”: Social policing of warlord government in Manchuria, 1918-1928
Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms
-
Decentralized multi-user multi-armed bandits with user dependent reward distributions
… is studied using a multi-player multi-armed bandits framework. We consider a decentralized multi-player stochastic multi-armed bandit model where the players cannot communicate with each other and can observe only their own actions and rewards. Furthermore, the environment may appear …
-
Dealers, insiders and bandits : learning and its effects on market outcomes
This thesis seeks to contribute to the understanding of markets populated by boundedly rational agents who learn from experience. Bounded rationality and learning have both been the focus of much research in computer science, economics and finance theory. However, we are at a critical stage in …
-
Bayesian tuning and bandits : an extensible, open source library for AutoML
The goal of this thesis is to build an extensible and open source library that handles the problems of tuning the hyperparameters of a machine learning pipeline, selecting between multiple pipelines, and recommending a pipeline. We devise a library that users can integrate into their existing …
-
NON-STATIONARY MULTIARMED BANDITS FOR SATIATION AND SEASONALITY PHENOMENA IN MUSIC RECOMMENDER SYSTEMS
… theoretical aspects of non-stationary multiarmed bandits, motivated by their application to music recommender systems. An intrinsic challenge of such systems lies in evolving user preferences. Rather than finding a single optimal item, the objective is to craft an ordered sequence of items …
-
Adaptive Preference Learning With Bandit Feedback: Information Filtering, Dueling Bandits and Incentivizing Exploration
… to existing work on classical multi-armed bandits, dueling bandits, and incentivizing exploration. For each type of feedback and application setting, we provide an algorithm and a theoretical analysis bounding its regret. We demonstrate through numerical experiments that our algorithms …
-
Theory and Practice of Large-scale Logistics: Offline Contextual Bandits and Decomposition Methods
… The first part focuses on offline contextual bandits, drawing from work at Uber on driver incentive programs. I introduce Empirical Soft Regret (ESR), a novel loss function for value-based learning that addresses limitations of accuracy-based approaches in misspecified settings. Unlike …
Page 1 of 4