Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 18 of 18 for “"Thompson sampling"”.
-
Particle Thompson sampling
Thompson sampling is an effective Bayesian heuristic for solving stochastic bandit problems. But it is hard to implement in practice due to the intractability of maintaining a continuous posterior distribution. Particle Thompson sampling (PTS) is an approximation of Thompson sampling based on the …
-
Low-complexity, low-regret link rate selection in rapidly varying wireless channels
… we consider a well-known algorithm called Thompson sampling to address this problem. However, unlike the traditional multi-armed bandit problem, a direct application of Thompson sampling results in a computational and storage complexity that grows exponentially with time. Therefore, we …
-
Improved worst-case regret bounds for randomized least-squares value iteration
… we introduce a clipping variant of one classical Thompson Sampling (TS)-like algorithm, randomized least-squares value iteration (RLSVI). Our $\tilde{\mathrm{O}}(H^2S\sqrt{AT})$ high-probability worst-case regret bound improves the previous sharpest worst-case regret bounds for RLSVI and matches …
-
Contextual Bandits with Neural Networks and Trees
… των ε-greedy, Upper Confidence Bound (UCB) και Thompson Sampling, οι οποίοι χρησιμοποιούνται για τη μείωση της μεταμέλειας (regret). Στη συνέχεια, η εργασία ανασκοπεί δύο από τις σημαντικότερες μεθό- δους στατιστικής μάθησης που χρησιμοποιούνται για τη μοντελοποίηση των σχέσεων μεταξύ πλαισίου …
-
Models of intelligence operations
… heuristics including knowledge gradient and Thompson sampling methods. The allocation policy is handled by thresholds which act as Lagrangian multipliers of the original MABA problem. Both a discrete Dirichlet-Multinomial and a continuous Exponential-Gamma-Gamma implementation of the MABA …
-
Robust sequential decision-making on networks
… Next, I developed a variant of the Bayesian Thompson Sampling algorithm in this setting, titled Robust [alpha]-TS, which involved developing an efficient pipeline for posterior inference. I also proved finite-time regret bounds for this algorithm, that are optimal up to logarithmic factors. …
-
Data-Driven Operations in Changing Environments
… that learns this prior online while solving a Thompson sampling pricing experiments for each product. Finally, motivated by our collaboration with AB InBev, a consumer packaged goods (CPG) company, we consider the problem of forecasting sales under the coronavirus disease 2019 (COVID-19) …
-
ONLINE RESOURCE ALLOCATION AND ITS APPLICATIONS
… on the reward distributions. We develop Thompson Sampling-style algorithms for mean-variance and CVaR MAB, and provide comprehensive regret analyses. Our algorithms achieve the best-known regret bounds for risk-aware MABs and also attain the information-theoretic bounds in some parameter …
-
Optimizing deep learning networks using multi-armed bandits
… (ii) Upper Confidence Bounds (UCB) (iii) Thompson Sampling and (iv) Exponential Weight Algorithm for Exploration and Exploitation (EXP3). The algorithms were implemented in Python and a comprehensive empirical evaluation of their performance was carried out in comparison to both the …
-
On upper confidence bound algorithms for piecewise-stationary stochastic multi-armed bandits and the variants
… (e.g., upper confidence bound (UCB) and Thompson sampling (TS)) have been proposed in the literature, which are order optimal compared with the lower bound. Original MAB problems are considered in a stationary environment, where the reward distributions do not evolve over time. Many …
-
Dynamic learning and optimization for operations management problems
… and pricing algorithm, which builds upon the Thompson sampling algorithm used for multi-armed bandit problems by incorporating inventory constraints. Our algorithm proves to have both strong theoretical performance guarantees as well as promising numerical performance results when compared to …
-
Analytics for online markets
… well-known multi-armed bandit algorithm called Thompson Sampling to consider a retailer's limited inventory constraints. Our algorithm has promising numerical performance results when compared to other algorithms developed for the same setting.
-
Multi-server federated learning in vehicular edge computing
… quality, cost, and availability. T-BIDS uses Thompson sampling to estimate each client’s contribution quality from validation improvements, and in each round selects participants by a quality-to-bid ratio under a per-round budget, with incentive-compatible payment rules. Clients can also …
-
Experimentation and Control in Online Platforms
… experimentation, dubbed Synthetically Controlled Thompson Sampling (SCTS), which robustly identifies the optimal treatment without the non-stationarity assumptions of the status quo, and minimizes the cost of experimentation by incurring near-optimal, square-root regret.
-
Combining artificial intelligence and robotic system in chemical product/process design
… classification algorithm with the Thompson-Sampling Efficient Multi-Optimization (TSEMO) for the simultaneous optimization of continuous and discrete outputs. The methodology was successfully applied to the design of a formulated liquid product of commercial interest for which no …
-
Decision-Making Under Uncertainty: From Theory to Practice
… two prominent policies for multi-armed bandits, Thompson sampling and upper confidence bound. We show that TS-UCB achieves materially lower regret on a comprehensive suite of synthetic and real-world datasets, and we establish optimal regret guarantees for TS-UCB for both the K-armed and linear …
-
Carbon Accounting for Sustainable Computing in Cloud Provisioned Data Centers
… using a multi-armed bandit algorithm using Thompson Sampling. This work illustrates that embodied emissions can constitute anywhere from five to thirty percent of a server's total environmental impact. Workloads can be more than 11 times higher between workloads given the same set of …
-
Use of Reinforcement Learning for Interference Avoidance or Efficient Jamming in Wireless Communications
… networks. This work explores the use of linear Thompson Sampling (TS) to jam OFDM-modulated signals. The jammer may select from both time-domain and frequency-domain jamming schemes. We demonstrate that the linear TS algorithm is able to perform better than a traditional reinforcement learning …