Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 15 of 15 for “"Contextual Bandits"”.
-
Contextual Bandits with Neural Networks and Trees
… με την εκμετάλλευση (exploitation). Τα bandits παρέχουν ένα απλό μοντέλο για αυτό το δίλημμα. Τα contextual bandits αποτελούν μια πολύ σημαντική κατηγορία, όπου ο πράκτορας(agent) έχει πρόσβαση σε πρόσθετες πληροφορίες που μπορεί να βοηθήσουν στην πρόβλεψη της ποιότητας των ενεργειών …
-
Theory and Practice of Large-scale Logistics: Offline Contextual Bandits and Decomposition Methods
… settings. The first part focuses on offline contextual bandits, drawing from work at Uber on driver incentive programs. I introduce Empirical Soft Regret (ESR), a novel loss function for value-based learning that addresses limitations of accuracy-based approaches in misspecified settings. …
-
Data-Driven Dynamic Decision Making: Algorithms, Structures, and Complexity Analysis
… effective dynamic decision making. Focusing on contextual bandits, a core class of online decision-making problems, we present the first optimal and efficient reduction from contextual bandits to offline regression. A remarkable consequence of our results is that advances in offline regression …
-
Fundamental Limits of Learning for Generalizability, Data Resilience, and Resource Efficiency
… noise (Chapter 4); partial-feedback in standard contextual bandits (Chapter 5) and, as a first step towards more complex reinforcement learning settings, contextual bandits with non-stationary or adversarial rewards (Chapter 6). We investigate the impact of resource constraints in Part II, …
-
Bandit Problems under Censored Feedback
… of problems through the lens of multi-armed and contextual bandits evolving in censored environments. Our goal is to estimate the performance loss due to censorship in the context of classical algorithms designed for uncensored environments. Our main contributions include the introduction of a …
-
Synthetic Data Generation and Sampling for Online Training of DNN in Manufacturing Supervised Learning Problems
… In the SIDES framework, a bi-level Hierarchical Contextual Bandits is proposed to provide a scientific way to integrate DoE and observational data sampling, which optimizes DNNs' online learning performance. Multimodality-aligned variational Autoencoder transforms the multimodal predictors from …
-
Trustworthy Machine Learning: From Algorithmic Transparency to Decision Support
… information. Using techniques from stochastic contextual bandits, we introduce THREAD, an online algorithm to personalize a decision support policy for each decision-maker. We deploy THREAD with real users to show how personalized policies can be learned online, and illustrate nuances of …
-
Adaptive optimization problems under uncertainty with limited feedback
… the set is strongly curved. Third, we study the Bandits with Knapsacks framework, a recent extension to the standard Multi-Armed Bandit framework capturing resource consumption. We extend the methodology developed for the original problem and design algorithms with regret bounds that are …
-
Enriched Story Experiences with a New Video Interaction Model
… module is achieved with the online learning of contextual bandits, using domain-specific features. Conversely, a story guidance module performs implicit emotion recognition on the user, as they engage with conversational agents representing story characters. Powered by natural language …
-
Anticipating, Extracting, and Leveraging Information in Clinical Decision-Making
… tasks: As the first example, we develop inverse contextual bandits (ICB), a method for learning how behavior evolves over time, and show that ICB can help assess the impact of new medical guidelines on actual clinical practice. As the second example, we define the notion of expertise, an …
-
Data Exchange for Artificial Intelligence Incubation in Manufacturing Industrial Internet
… (Chapter 2); (2) an ensemble active learning by contextual bandits framework for acquisition and evaluation of passively collected online data for the continuous improvement and resilient modeling performance during the online training and deployment phases (Chapter 3); and (3) a context-aware, …
-
Decision-Making Under Uncertainty: From Theory to Practice
… our algorithmic framework with a case study on contextual bandits for warfarin dosing where we are concerned with the cost of exploration across multiple races and age groups. Next, we study the classical problem of minimizing regret for multi-armed bandits. In this classic problem, there are …
-
Use of Reinforcement Learning for Interference Avoidance or Efficient Jamming in Wireless Communications
… a reinforcement learning algorithm called contextual bandits. The harsh environment of an underwater channel provides a challenging problem. The channel may induce multipath and time delays which lead to time-varying, frequency-selective attenuation. These factors are also influenced by the …
-
Principled exploration in sequential decision-making
Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-05-01
-
Coactive Learning Algorithms for Constructive Preference Elicitation
… approaches, such as collaborative filtering and contextual bandits, while in the latter case data is usually scarce, making it necessary to employ specialized algorithms for preference elicitation. Preference elicitation algorithms interactively build a utility model of the user preferences and …