What is Importance Sampling?
Importance Sampling is a change-of-measure method that estimates an expectation under a target distribution by reweighting samples drawn from a different proposal distribution with the density ratio between target and proposal.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Choose the proposal around the estimand
A proposal should cover every material target region and place enough probability where |f(x)|p(x) is large. A proposal that merely resembles the target may still be poor for a rare-event estimand. Record whether target and proposal densities are normalized, compute ratios in log space when densities are tiny, and treat any non-finite ratio or target-positive/proposal-zero event as a support failure rather than a row to discard.
Ordinary and self-normalized estimators answer slightly different questions
Ordinary IS is unbiased when the normalized target density is known and the weighted integrand is integrable. Self-normalized IS can use an unnormalized target and often behaves more stably, but it is generally biased at finite sample sizes. Clipping, truncation, or stabilizing weights can further reduce variance while changing bias and sometimes the effective target; those choices belong in the estimator contract.
Diagnose weights before interpreting the estimate
Report raw sample count alongside Effective Sample Size, maximum normalized weight, upper weight quantiles, repeated-run variability, and sensitivity to defensible clipping thresholds. Pareto Smoothed Importance Sampling uses fitted tail behavior to diagnose and stabilize problematic weights, but no smoothing technique creates missing support. When diagnostics collapse, collect better proposal data or narrow the target claim.
Key Characteristics
- Reweights proposal samples by a target-to-proposal density ratio
- Requires target support to be contained in proposal support
- Supports ordinary and self-normalized estimators
- Can have extreme variance when target and proposal tails mismatch
- Uses Effective Sample Size and tail diagnostics to expose concentration
- Trades bias against variance when weights are clipped or stabilized
Common Use Cases
- Estimating rare-event probabilities from a targeted simulation proposal
- Correcting evaluation data toward a declared deployment distribution
- Evaluating a target decision policy from randomized logging traffic
- Reusing posterior samples after a moderate model or prior change
- Comparing proposal designs before an expensive Monte Carlo run
Example
Loading code...Frequently Asked Questions
When is an Importance Sampling estimate unbiased?
Ordinary Importance Sampling is unbiased when samples truly come from the proposal, the target density is normalized, target support is covered by proposal support, and the weighted quantity is integrable. Estimated densities, adaptive reuse, dependence, clipping, or self-normalization require separate bias and uncertainty analysis.
What is the difference between ordinary and self-normalized Importance Sampling?
Ordinary IS divides the weighted sum by the number of proposal samples and needs the normalized density ratio. Self-normalized IS divides by the observed sum of weights, can use an unnormalized target, and is generally biased at finite sample sizes but often more stable.
Why can Importance Sampling fail even with many samples?
Raw row count does not ensure information. If a few tail observations carry nearly all weight, variance may be enormous and the Effective Sample Size may remain tiny as more ordinary samples are added. Missing proposal support is more severe because the target expectation is not identified from those samples.
Should Importance Sampling weights be clipped?
Clipping may lower variance and limit operational instability, but it introduces bias and can change the population being approximated. Report the rule, compare justified thresholds, retain unclipped diagnostics, and do not use clipping to conceal a support violation.
Is Importance Sampling the same as Inverse Propensity Scoring?
Inverse Propensity Scoring is an Importance Sampling construction specialized to decisions or treatments. Its weights use target and logging action probabilities conditional on context, while generic Importance Sampling can reweight arbitrary probability distributions and integrands.