What is Thompson Sampling?
Thompson Sampling is a randomized sequential decision policy that samples one plausible model from the current posterior and chooses the action that is optimal under that sampled model.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Explore in proportion to posterior plausibility
A model is sampled from the posterior, then the action optimal for that model is chosen. Consequently, an action is selected with the posterior probability that it is optimal under exact probability matching. Russo and colleagues' tutorial develops this rule from Bernoulli bandits to structured decisions and discusses prior specification, approximations, nonstationarity, constraints, and settings where simple posterior sampling can fail.
Model feedback, context, and operational constraints
Beta-Bernoulli updates apply only to binary, conditionally independent rewards with stationary arm probabilities. Continuous, contextual, censored, delayed, or drifting outcomes need an appropriate likelihood and posterior approximation. Selection changes which outcomes are observed, so preserve action propensities, context, assignment time, reward window, missing feedback, and model version. Safety or policy constraints must filter or penalize actions before random exploration reaches users.
Evaluate cumulative decisions and posterior sensitivity
Bandit policies target cumulative reward or regret, not only identification of the best arm after a fixed experiment. Compare cumulative and per-round regret, reward, unsafe-action rate, subgroup exposure, switching cost, posterior concentration, and recovery after drift across repeated simulations and online guardrails. Agrawal and Goyal prove guarantees for a specific linear contextual-bandit formulation; those bounds do not transfer unchanged to arbitrary models or approximate posteriors.
Key Characteristics
- Samples a plausible model or parameter from the current posterior
- Acts optimally for that sample instead of adding a fixed exploration bonus
- Randomizes exploration according to modeled uncertainty
- Supports independent, contextual, structured, and function-space decisions
- Depends on posterior quality, feedback timing, and stationarity assumptions
- Is evaluated by cumulative reward or regret under explicit constraints
Common Use Cases
- Allocating traffic among online product variants
- Choosing recommendations with contextual reward models
- Selecting experiments through posterior function sampling
- Exploring treatment or policy alternatives under strict safeguards
- Adapting resource allocation while outcomes arrive sequentially
Example
Loading code...Frequently Asked Questions
How does Thompson Sampling balance exploration and exploitation?
It samples from the posterior and chooses the action optimal for that sampled world. Well-supported actions win often, while uncertain actions still win when a plausible posterior draw makes them best. Exploration therefore follows modeled uncertainty rather than a fixed random rate.
Is Thompson Sampling the same as A/B testing?
No. A fixed-horizon randomized A/B test is usually designed for estimation or hypothesis testing under a prespecified allocation. Thompson Sampling adapts allocation to observed outcomes to optimize cumulative decisions, which changes exposure and requires sequential-analysis-aware inference.
Does Thompson Sampling require a Beta distribution?
No. Beta-Bernoulli is a convenient conjugate example for binary rewards. Gaussian, generalized-linear, neural, Gaussian Process, or other posterior models can be used when they match the reward and context, often with approximate inference.
How should delayed rewards be handled in Thompson Sampling?
Track pending actions separately and update only after their attribution window closes. Ignoring delay can oversample an arm because its failures or successes are missing. Simulate the real delay process and define how late, censored, and duplicate outcomes are reconciled.
When can Thompson Sampling perform poorly?
It can fail under misspecified or overconfident posteriors, nonstationary rewards, weak priors with scarce data, unsafe exploration, unmodeled context, delayed feedback, or objectives where information has complex long-horizon value. Guardrails and alternative baselines remain necessary.