Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 8 of 8 for “"bandit feedback"”.
-
Adaptive Preference Learning With Bandit Feedback: Information Filtering, Dueling Bandits and Incentivizing Exploration
… learning system learns users' preferences from feedback while simultaneously using these learned preferences to help them find preferred items. We study three different types of user feedback in three application setting: cardinal feedback with application in information filtering systems, …
-
Multi-armed bandits and applications to large datasets
This thesis considers the multi-armed bandit (MAB) problem, both the traditional bandit feedback and graphical bandits when there is side information. Motivated by the Boltzmann exploration algorithm often used in the more general context of reinforcement learning, we present Almost Boltzmann …
-
Online and active learning of big networks: theory and algorithms
… active learning), and online learning with bandit feedback algorithms for learning in a network. In the first part of this thesis, I propose a \textit{nonadaptive} active learning approach on a graph, based on generalization error bound minimization. In particular, I present a data-dependent …
-
Sequential decision making with feature-linear models
… as linear in some feature space. We consider a bandit problem, where the rewards are linear in a reproducing kernel Hilbert space, and a reinforcement learning setting with features given by a neural network. The thesis is split into two parts accordingly. In part I, we introduce a new algorithm …
-
LEARNING UNDER STRUCTURE AND UNCERTAINTY: ALGORITHMS FOR BANDIT AND ONLINE DECISION MAKING
… with a particular focus on the multi-armed bandit and online learning frameworks. At the core of these settings lies a simple yet powerful interaction protocol: at each round the learner receives a question (information available prior to a decision), provides an answer (an action or …
-
New Directions in Bandit Learning: Singularities and Random Walk Feedback
<p>My thesis focuses new directions in bandit learning problems. In Chapter 1, I give an overview of the bandit learning literature, which lays the discussion framework for studies in Chapters 2 and 3. In Chapter 2, I study bandit learning problem in metric measure spaces. I start with multi-armed …
-
Online Combinatorial Optimization for Digital Marketplaces
… is applicable in both full-information and bandit feedback settings, obtaining $\sqrt{T}$ and $T^{3/4}$ regret respectively. In the second chapter, we focus on the problem of learning non-parametric choice models on digital platforms in an active learning setting. This method involves …
-
Learning-NUM: Utility Maximization in Stochastic Queueing Networks
… of learning and network control and dealing with feedback delay and unknown constraints. We start by considering Learning-NUM problems with linear utility functions in bipartite networks, where the corresponding static optimization problems are linear programs. We propose a priority-based network …