What is Off-Policy Evaluation?
Off-Policy Evaluation is the estimation of a target policy's expected value using logged outcomes collected by a different behavior or logging policy, without first deploying the target policy.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Identification begins with consistency and overlap
The logged reward must correspond to the recorded action under a stable outcome definition, and the logging probability must describe how that action was selected from the eligible set. Every action the target policy may choose needs positive logging probability in the relevant context. Hidden overrides, missing candidates, interference, unrecorded confounding, and post-decision features can invalidate the estimate before an estimator is chosen.
Estimator choice is a bias-variance decision
IPS is unbiased under correct propensities and support but can have extreme variance; Direct Method can generalize to rare actions but inherits reward-model bias. Dudik, Langford, and Li adapt Doubly Robust estimation to contextual policy evaluation: its correction can remain unbiased when either the reward model or propensity model is correct under the remaining assumptions. Weight clipping, self-normalization, and model reuse change finite-sample behavior and must be disclosed.
Evaluate the evaluator before trusting a policy ranking
Freeze target policies before the final test, use cross-fitting where nuisance models could overfit, report confidence intervals, effective sample size, maximum and tail weights, support violations, and results by important slice. Compare estimators through simulation or randomized ground truth before production use. The Open Bandit Dataset and Pipeline was created to make such real-world estimator comparisons reproducible; one favorable benchmark still does not certify a new deployment distribution.
Key Characteristics
- Uses outcomes collected by a behavior policy to evaluate another policy
- Requires decision-time context, action, reward, and logging propensity
- Depends on consistency, support, and correct timing assumptions
- Includes Direct Method, IPS, SNIPS, and Doubly Robust estimators
- Trades reward-model bias against importance-weight variance
- Must report uncertainty and weight diagnostics, not only a point estimate
Common Use Cases
- Screening recommendation policies before an online experiment
- Comparing contextual-bandit policies from randomized traffic logs
- Estimating a routing policy's quality-cost trade-off offline
- Auditing treatment policies when exploration probabilities are recorded
- Selecting safer candidates for a controlled canary deployment
Example
Loading code...Frequently Asked Questions
What data is required for Off-Policy Evaluation?
A contextual-bandit record usually needs decision-time context, the eligible action set, chosen action, its probability under the logging policy, observed reward and attribution window, plus policy and feature versions. Sequential reinforcement-learning OPE additionally needs ordered transitions, termination, and horizon semantics.
Why does Off-Policy Evaluation need overlap?
A log can only identify outcomes for actions the logging policy had a chance to select. If the target policy chooses an action with zero logging probability in some context, no importance weight or reward correction can reconstruct that missing evidence without additional modeling assumptions.
Why is Doubly Robust estimation called doubly robust?
Under its full identification conditions, the estimator can remain consistent when either the reward model is correct or the propensity model is correct, rather than requiring both nuisance models to be correct. It is not immune to support violations, bad outcome timing, dependence, or both models being wrong.
Can a deterministic logging policy support OPE?
Only for target decisions already supported by that deterministic policy, unless defensible structural assumptions or external randomized data provide identification. Estimating propensities after the fact does not create overlap where the logging system never selected an action.
Can Off-Policy Evaluation replace an online A/B test?
Usually it should screen and de-risk candidates, not serve as unconditional replacement. OPE is sensitive to overlap, model error, logging defects, and distribution shift. High-impact changes still need a controlled rollout or experiment with safety gates and rollback.