Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 19 of 19 for “"Thompson sampling"”.
-
Particle Thompson sampling
Thompson sampling is an effective Bayesian heuristic for solving stochastic bandit problems. But it is hard to implement in practice due to the intractability of maintaining a continuous posterior distribution. Particle Thompson sampling (PTS) is an approximation of Thompson sampling based on the …
-
Low-complexity, low-regret link rate selection in rapidly varying wireless channels
… we consider a well-known algorithm called Thompson sampling to address this problem. However, unlike the traditional multi-armed bandit problem, a direct application of Thompson sampling results in a computational and storage complexity that grows exponentially with time. Therefore, we …
-
Improved worst-case regret bounds for randomized least-squares value iteration
… we introduce a clipping variant of one classical Thompson Sampling (TS)-like algorithm, randomized least-squares value iteration (RLSVI). Our $\tilde{\mathrm{O}}(H^2S\sqrt{AT})$ high-probability worst-case regret bound improves the previous sharpest worst-case regret bounds for RLSVI and matches …
-
Contextual Bandits with Neural Networks and Trees
… των ε-greedy, Upper Confidence Bound (UCB) και Thompson Sampling, οι οποίοι χρησιμοποιούνται για τη μείωση της μεταμέλειας (regret). Στη συνέχεια, η εργασία ανασκοπεί δύο από τις σημαντικότερες μεθό- δους στατιστικής μάθησης που χρησιμοποιούνται για τη μοντελοποίηση των σχέσεων μεταξύ πλαισίου …
-
Models of intelligence operations
… heuristics including knowledge gradient and Thompson sampling methods. The allocation policy is handled by thresholds which act as Lagrangian multipliers of the original MABA problem. Both a discrete Dirichlet-Multinomial and a continuous Exponential-Gamma-Gamma implementation of the MABA …
-
Robust sequential decision-making on networks
… Next, I developed a variant of the Bayesian Thompson Sampling algorithm in this setting, titled Robust [alpha]-TS, which involved developing an efficient pipeline for posterior inference. I also proved finite-time regret bounds for this algorithm, that are optimal up to logarithmic factors. …
-
Data-Driven Operations in Changing Environments
… that learns this prior online while solving a Thompson sampling pricing experiments for each product. Finally, motivated by our collaboration with AB InBev, a consumer packaged goods (CPG) company, we consider the problem of forecasting sales under the coronavirus disease 2019 (COVID-19) …
-
ONLINE RESOURCE ALLOCATION AND ITS APPLICATIONS
… on the reward distributions. We develop Thompson Sampling-style algorithms for mean-variance and CVaR MAB, and provide comprehensive regret analyses. Our algorithms achieve the best-known regret bounds for risk-aware MABs and also attain the information-theoretic bounds in some parameter …
-
Probabilistic user modelling methods for improving human-in-the-loop machine learning for prediction
… Sequential inference on the joint model, using Thompson sampling, was employed to find the targeted recommendation with minimum interaction. Simulated experiments and user studies in both tasks demonstrate improved prediction performance only after few interactions with the users. The research …
-
Optimizing deep learning networks using multi-armed bandits
… (ii) Upper Confidence Bounds (UCB) (iii) Thompson Sampling and (iv) Exponential Weight Algorithm for Exploration and Exploitation (EXP3). The algorithms were implemented in Python and a comprehensive empirical evaluation of their performance was carried out in comparison to both the …
-
On upper confidence bound algorithms for piecewise-stationary stochastic multi-armed bandits and the variants
… (e.g., upper confidence bound (UCB) and Thompson sampling (TS)) have been proposed in the literature, which are order optimal compared with the lower bound. Original MAB problems are considered in a stationary environment, where the reward distributions do not evolve over time. Many …
-
Dynamic learning and optimization for operations management problems
… and pricing algorithm, which builds upon the Thompson sampling algorithm used for multi-armed bandit problems by incorporating inventory constraints. Our algorithm proves to have both strong theoretical performance guarantees as well as promising numerical performance results when compared to …
-
Analytics for online markets
… well-known multi-armed bandit algorithm called Thompson Sampling to consider a retailer's limited inventory constraints. Our algorithm has promising numerical performance results when compared to other algorithms developed for the same setting.
-
Multi-server federated learning in vehicular edge computing
… quality, cost, and availability. T-BIDS uses Thompson sampling to estimate each client’s contribution quality from validation improvements, and in each round selects participants by a quality-to-bid ratio under a per-round budget, with incentive-compatible payment rules. Clients can also …
-
Experimentation and Control in Online Platforms
… experimentation, dubbed Synthetically Controlled Thompson Sampling (SCTS), which robustly identifies the optimal treatment without the non-stationarity assumptions of the status quo, and minimizes the cost of experimentation by incurring near-optimal, square-root regret.
-
Combining artificial intelligence and robotic system in chemical product/process design
… classification algorithm with the Thompson-Sampling Efficient Multi-Optimization (TSEMO) for the simultaneous optimization of continuous and discrete outputs. The methodology was successfully applied to the design of a formulated liquid product of commercial interest for which no …
-
Decision-Making Under Uncertainty: From Theory to Practice
… two prominent policies for multi-armed bandits, Thompson sampling and upper confidence bound. We show that TS-UCB achieves materially lower regret on a comprehensive suite of synthetic and real-world datasets, and we establish optimal regret guarantees for TS-UCB for both the K-armed and linear …
-
Carbon Accounting for Sustainable Computing in Cloud Provisioned Data Centers
… using a multi-armed bandit algorithm using Thompson Sampling. This work illustrates that embodied emissions can constitute anywhere from five to thirty percent of a server's total environmental impact. Workloads can be more than 11 times higher between workloads given the same set of …
-
Use of Reinforcement Learning for Interference Avoidance or Efficient Jamming in Wireless Communications
… networks. This work explores the use of linear Thompson Sampling (TS) to jam OFDM-modulated signals. The jammer may select from both time-domain and frequency-domain jamming schemes. We demonstrate that the linear TS algorithm is able to perform better than a traditional reinforcement learning …