Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 16 of 16 for “"Regret minimization"”.
-
Computing Equilibria in Colonel Blotto by Applying Counterfactual Regret Minimization Using a Layered Graph Representation
… from [3] and use it to run counterfactual regret minimization (CFR) on Colonel Blotto for the first time. CFR is a state-of-the-art learning algorithm that permits parameter free learning and has practical performance that is much better than its theoretical bounds.
-
Improved worst-case regret bounds for randomized least-squares value iteration
This work studies regret minimization with randomized value functions in reinforcement learning. In tabular finite-horizon Markov Decision Processes, we introduce a clipping variant of one classical Thompson Sampling (TS)-like algorithm, randomized least-squares value iteration (RLSVI). Our …
-
New Models qnd Algorithms for Bandits and Markets
… a generic structural assumption on rewards to regret rates in an online optimization problem, is not fully developed. The primary goal of this dissertation, therefore, will be to fill out the space of models, algorithms, and assumptions used in sequential decision making problems. Toward this …
-
On Solving Larger Games: Designing New Algorithms Adaptable to Deep Reinforcement Learning
… Chapter 4 shows that algorithms based on regret decomposition enjoy best-iterate convergence to the NE. Chapter 5 proposes Q-value based Regret Minimization (QFR), which achieves all three properties simultaneously.
-
Online Auctions with Multiple Items
… parallel. In particular, we study the problem of regret minimization in this setting, extending prior work for second-price auctions. We show that sub-linear regret cannot be achieved when the values are continuous and there are two or more single-item auctions that take place per round. On the …
-
Efficient Learning and Computation of Linear Correlated Equilibrium in General Convex Games
We propose efficient no-regret learning dynamics and ellipsoid-based methods for computing linear correlated equilibria—a relaxation of correlated equilibria and a strengthening of coarse correlated equilibria—in general convex games. These are games where the number of pure strategies is …
-
Multiagent Approaches to Enhance Learning and Trust in AI Systems
… learning, this work demonstrates how regret minimization can serve as a unifying foundation for scalable learning. Contributions include a new interpretation of an existing Multiagent Reinforcement Learning (MARL) as a regret-based method, improvements its learning through modified …
-
Learning Effective and Human-like Policies for Strategic, Multi-Agent Games
… a more general class of policies. We develop a regret-minimization algorithm for imperfect information games that can leverage human demonstrations. We show that using this algorithm for search in no-press Diplomacy yields a policy that matches the human-likeness of IL while achieving much …
-
Advanced Ordered Weighted Averaging Methods in Robust Optimization
… approaches, such as min-max and min-max regret, focus on minimizing the worst-case outcomes and worst-case regret, respectively, often resulting in highly conservative solutions. To address this limitation, this dissertation investigates the Ordered Weighted Averaging (OWA) operator, …
-
Building Strategic AI Agents for Human-centric Multi-agent Systems
… through two approaches. We start by developing a regret minimization algorithm for modeling actions of strong and human-like agents called piKL, which incorporates a cost term proportional to the KL divergence between a search policy and a humanimitation learned policy. This approach improves …
-
Mode and Departure Time Choice Behavior of Non-Work Related Trips
… part analyzes departure time choice. A Random Regret-Minimization (RRM) approach is applied on stated preference data that was collected in Calgary, Canada. We study mode choice behavior by examining the impacts of various sustainable neighborhood design elements like availability of …
-
Online Learning in Control Theory
… The controller aims to achieve a sublinear regret, similar to the online optimization setting. A modified online Riccati algorithm is introduced that under some uniform boundedness assumptions results in a logarithmic regret bound. In particular, the logarithmic regret for the scalar case is …
-
Low-complexity, low-regret link rate selection in rapidly varying wireless channels
… states and which achieves at most logarithmic regret as a function of time when compared to an optimal algorithm which knows the probability distribution of the channel states.
-
RISK- AND AMBIGUITY-AVERSE OPTIMIZATION MADE MORE TRACTABLE AND LESS CONSERVATIVE
… exhibits risk aversion, ambiguity aversion and regret aversion. Theoretically, we develop a new regret-based robust satisficing framework and devise an upper confidence bound algorithm to tackle two-stage optimization and online learning, respectively. The thesis is complemented by several …
-
Robustness of the k-double auction under Knightian uncertainty
… auctions when those markets are populated by regret minimizers. Regret minimizing agents, unlike typical expected utility maximizers, need not commit to a single prior in their decision rule. In fact, it is a feature of the minimax regret decision rule that is not based on any prior. This …
-
Leveraging Structured Feedback in Online Learning
L'abstract è presente nell'allegato / the abstract is in the attachment