What is Contextual Bandit?

A Contextual Bandit is a sequential decision problem in which a policy observes the current context, chooses an available action, receives only that action's reward, and learns to maximize reward across future contexts.

Quick Facts

SpecificationOfficial Specification

How It Works

Use context without inventing counterfactual labels

At each round, the system observes context and available actions before selecting one. Only the selected action's reward becomes observable. Li, Chu, Langford, and Schapire formalized personalized news recommendation this way and introduced LinUCB for a linear payoff assumption. Their replay evaluation relies on randomized logged traffic; it does not make arbitrary production logs unbiased.

Separate reward modeling from exploration

A predictor estimates action value, while an exploration policy determines which evidence will be collected next. LinUCB adds a confidence bonus, Thompson Sampling samples plausible parameters, and policy-class methods optimize broader decision rules. Agarwal and colleagues show how general contextual-bandit guarantees depend on partial feedback, a comparator policy class, and controlled exploration, not merely a high-accuracy supervised model.

Log the decision contract and test support

Store context as seen at decision time, eligible actions, chosen action, its probability under the logging policy, model and feature versions, reward window, missing outcomes, and overrides. Before offline evaluation, verify that the logging policy assigned positive probability to actions the target policy may choose. Monitor cumulative reward, regret proxies, propensity and weight distributions, subgroup exposure, safety violations, latency, and behavior under delayed or drifting feedback.

Key Characteristics

  • Conditions each decision on features observed before action selection
  • Reveals reward only for the selected action
  • Learns a policy that maps contexts to action distributions
  • Combines a value model with an explicit exploration mechanism
  • Requires propensity and eligibility logs for counterfactual evaluation
  • Remains a one-step model unless action-dependent state transitions are added

Common Use Cases

  1. Personalizing a recommendation or message for the current request
  2. Selecting an offer while learning heterogeneous treatment response
  3. Routing prompts among models using quality, cost, and latency context
  4. Choosing just-in-time interventions under safety constraints
  5. Adapting interface variants for different user or session conditions

Example

loading...
Loading code...

Frequently Asked Questions

What is the difference between a Multi-Armed Bandit and a Contextual Bandit?

A basic Multi-Armed Bandit learns one reward distribution per arm and seeks the best overall action. A Contextual Bandit observes features before acting and learns which action works for each context. Both reveal only the chosen action's reward.

Is a Contextual Bandit just a recommendation model?

No. A recommendation model can rank items from historical labels without controlling data collection. A contextual bandit includes an action policy, exploration, partial feedback, and a cumulative decision objective. Recommendation is one application, and production recommenders often contain many non-bandit stages.

Does LinUCB work with any contextual relationship?

No. Standard LinUCB assumes expected reward is linear in the chosen feature representation and uses a confidence construction tied to that model. Nonlinear effects, omitted features, nonstationarity, delayed outcomes, and invalid uncertainty estimates can break its behavior or theoretical guarantee.

What must be logged for offline Contextual Bandit evaluation?

At minimum, log the context available at decision time, eligible action set, chosen action, its logging-policy probability, observed reward and attribution window, policy version, and any overrides. Without reliable propensities and overlap, IPS or doubly robust estimates cannot recover unsupported actions.

When should I use reinforcement learning instead of a Contextual Bandit?

Use a sequential reinforcement-learning model when actions materially change future states, rewards depend on long trajectories, or credit assignment spans multiple decisions. A contextual bandit is appropriate when the main consequence can be treated as a one-step reward for the current context.

Related Terms

Related Articles