What is Inverse Propensity Scoring?

Inverse Propensity Scoring is an importance-weighting estimator that evaluates a target policy or treatment regime by multiplying each observed outcome by the ratio of its target action probability to its logging or assignment probability.

Quick Facts

SpecificationOfficial Specification

How It Works

Logging propensity is a decision-time probability

The denominator must be the probability that the deployed logger assigned the recorded action from the eligible set using information available at that decision. Post-action features, a probability from a newer policy version, or the model score before business-rule overrides is not equivalent. Persist candidate eligibility, chosen action, final probability, policy version, context version, reward window, and override reason in an immutable event.

Overlap determines what can be evaluated

If pi(a|x) > 0 while mu(a|x) = 0, the target policy asks about an action absent from the evidence. Very small but nonzero propensities produce extreme weights and high variance. Counterfactual Risk Minimization shows how propensity-weighted objectives support learning from logged bandit feedback, while also making variance control part of the optimization problem.

Point estimates need weight and uncertainty diagnostics

Report ordinary IPS, Self-Normalized IPS (SNIPS), Effective Sample Size, maximum and high-quantile weights, support violations, and results by consequential slice. SNIPS often lowers variance but is generally biased at finite sample sizes. Clipping also introduces bias. Predeclare thresholds, bootstrap or otherwise derive uncertainty with the correct dependence unit, and compare against a Direct Method or Doubly Robust estimate.

Key Characteristics

  • Uses target-to-logging action-probability ratios
  • Corrects selection bias only under explicit identification assumptions
  • Requires positive overlap for every target action being evaluated
  • Can be unbiased yet have unusably high variance
  • Includes ordinary IPS and Self-Normalized IPS variants
  • Needs immutable decision-time propensities and weight diagnostics

Common Use Cases

  1. Screening a recommendation policy before an online experiment
  2. Evaluating contextual-bandit candidates on exploration traffic
  3. Estimating a treatment regime from randomized assignment logs
  4. Reweighting moderation or routing decisions to a target policy
  5. Auditing how weak overlap changes a policy ranking

Example

loading...
Loading code...

Frequently Asked Questions

What exactly is a propensity in IPS?

For policy evaluation, it is the probability that the logging policy selected the recorded action given the decision-time context and eligible actions. In causal inference, it commonly denotes treatment probability given observed pre-treatment covariates. The probability must correspond to the actual assignment mechanism.

What is the difference between IPS and SNIPS?

Ordinary IPS averages weighted rewards over the raw number of records and can be unbiased under its assumptions. SNIPS divides by the observed sum of weights, which often stabilizes scale and variance but generally introduces finite-sample bias.

Can estimated propensities replace missing logging probabilities?

Only under stronger modeling and causal assumptions, with uncertainty that includes propensity estimation. A fitted model cannot recreate randomized evidence or positive overlap for actions the production system never selected, and post-treatment inputs can make the estimate invalid.

How should extreme IPS weights be handled?

First investigate eligibility, policy-version, and logging defects. Then report the original weight distribution and Effective Sample Size. Clipping or stabilization may be used as a declared bias-variance trade-off with sensitivity analysis, not as silent cleanup.

Does a stable IPS estimate prove a target policy is safe to deploy?

No. Stability does not prove consistency, absence of hidden confounding, correct reward attribution, or future distribution stability. IPS should narrow candidates for a controlled experiment or rollout with guardrails, monitoring, and rollback.

Related Terms