Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 14 of 14 for “"Bandit Problems"”.
-
Bandit Problems under Censored Feedback
… transportation. We model and study this class of problems through the lens of multi-armed and contextual bandits evolving in censored environments. Our goal is to estimate the performance loss due to censorship in the context of classical algorithms designed for uncensored environments. Our main …
-
Regulating exploration in multi-armed bandit problems with time patterns and dying arms
… holidays. The standard paradigm of multi-armed bandit analysis does not take these known patterns into account. This means that for applications in retail, where prices are fixed for periods of time, current bandit algorithms will not suffice. This work provides a framework and methods that take …
-
Particle Thompson sampling
… Bayesian heuristic for solving stochastic bandit problems. But it is hard to implement in practice due to the intractability of maintaining a continuous posterior distribution. Particle Thompson sampling (PTS) is an approximation of Thompson sampling based on the simple idea of replacing …
-
New Models qnd Algorithms for Bandits and Markets
… consider large-scale sequential decision making problems in which a learner must deploy an algorithm to behave optimally under uncertainty. Although many of these problems can be modeled as contextual bandit problems, we argue that the tools and techniques for analyzing bandit problems with large …
-
Bayesian Analysis, Endogenous Data,and Convergence of Beliefs
Problems in statistical analysis, economics, and many other disciplines often involve a trade-off between rewards and additional information that could yield higher future rewards. This thesis investigates such a trade-off, using a class of problems known as bandit problems. In these problems, a …
-
Online decision problems with large strategy sets
… We study an important class of online decision problems called generalized multi- armed bandit problems. In the past such problems have found applications in areas as diverse as statistics, computer science, economic theory, and medical decision-making. Most existing algorithms were efficient …
-
Dynamic learning and optimization for operations management problems
… Thompson sampling algorithm used for multi-armed bandit problems by incorporating inventory constraints. Our algorithm proves to have both strong theoretical performance guarantees as well as promising numerical performance results when compared to other algorithms developed for similar settings. …
-
Data, models and decisions for large-scale stochastic optimization problems
… this question in the context of four important problems: the dynamic control of large-scale stochastic systems, the design of product lines under uncertainty, the selection of an assortment from historical transaction data and the design of a personalized assortment policy from data. In the …
-
Enhancing Capabilities of Assistive Robotic Arms: Learning, Control, and Object Manipulation
… task learning as a sequence of multi-armed bandit problems, where each bandit problem corresponds to a waypoint in the robot's trajectory. We introduce an approximate posterior sampling solution that builds the robot's motion one waypoint at a time. Our simulations and real-world experiments …
-
Online advertisements and multi-armed bandits
We investigate a number of multi-armed bandit problems that model different aspects of online advertising, beginning with a survey of the key techniques that are commonly used to demonstrate the theoretical limitations and achievable results for the performance of multi-armed bandit algorithms. We …
-
NON-STATIONARY MULTIARMED BANDITS FOR SATIATION AND SEASONALITY PHENOMENA IN MUSIC RECOMMENDER SYSTEMS
… theoretical aspects of non-stationary multiarmed bandits, motivated by their application to music recommender systems. An intrinsic challenge of such systems lies in evolving user preferences. Rather than finding a single optimal item, the objective is to craft an ordered sequence of items …
-
New Spatio-temporal Hawkes Process Models For Social Good
… integrate Hawkes Process models with multi-armed bandit algorithms, high dimensional marks, and high-dimensional auxiliary data to solve problems in search and rescue, forecasting infectious disease, and early detection of overdose spikes. In Chapter 3, we develop a method applications to the …
-
Uncertainty in interactive decision-making: learning, incentives, and robustness
Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms
-
Analysis based on incomplete data
Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2024-12-01