What is Effective Sample Size?

Effective Sample Size is a diagnostic that expresses how much information a weighted or dependent sample contains relative to an idealized sample of equally weighted independent observations, using a formula matched to the sampling process.

Quick Facts

SpecificationOfficial Specification

How It Works

Weight ESS measures concentration, not truth

Because the ratio is invariant to a common rescaling of weights, it summarizes their relative inequality. It is useful for Importance Sampling, IPS, and Covariate Shift diagnostics, but it does not test whether weights are correct, whether the target has support, or whether outcomes are independent. A high ESS with systematically wrong weights can produce a precise-looking wrong answer.

Use the ESS definition that matches the dependence mechanism

Weight ESS and MCMC ESS address different losses of information. MCMC ESS uses the integrated autocorrelation structure of a specific estimand, while survey design effects may include clusters and strata. Elvira, Martino, and Robert show that the conventional Importance Sampling ESS is a useful approximation with important limitations; report the formula and context instead of writing only “ESS.”

Pair ESS with tail, slice, and estimator diagnostics

Report normalized ESS (ESS / n), maximum normalized weight, upper quantiles, zero-weight fraction, and ESS by consequential slice. Compare results under justified weight caps while disclosing the induced bias. For confidence intervals, use an estimator-specific variance, bootstrap, or asymptotic method that respects clusters, users, sessions, or time; substituting ESS into an independent-sample formula is not automatically valid.

Key Characteristics

  • Converts weight concentration or sample dependence into an equivalent count
  • Equals raw sample size for equal independent weighted observations
  • Uses `(sum w)^2 / sum(w^2)` for common weight diagnostics
  • Uses autocorrelation-based definitions for MCMC draws
  • Does not validate weights, support, independence, or model assumptions
  • Must be paired with tail diagnostics and estimator-specific uncertainty

Common Use Cases

  1. Diagnosing concentrated weights in Off-Policy Evaluation
  2. Comparing proposals in an Importance Sampling experiment
  3. Monitoring density-ratio weights under Covariate Shift
  4. Assessing autocorrelated posterior draws in MCMC
  5. Defining release gates for unstable weighted metrics

Example

loading...
Loading code...

Frequently Asked Questions

What does weight-based Effective Sample Size measure?

The common `(sum w)^2 / sum(w^2)` diagnostic measures how concentrated nonnegative weights are relative to equal weights. It approximates information loss from unequal weighting, but does not include outcome variation, dependence, model error, or support failure.

Is Effective Sample Size always smaller than the raw sample size?

The Kish weight formula is at most `n` for nonnegative weights and equals `n` when all weights are equal. Other definitions can behave differently; some autocorrelation estimates may exceed `n` under negative correlation, so the formula and estimator must be named.

What is the difference between weight ESS and MCMC ESS?

Weight ESS summarizes inequality in importance or survey weights. MCMC ESS estimates information loss from serial autocorrelation for a particular parameter or function. They diagnose different mechanisms and are not interchangeable.

What is a good Effective Sample Size threshold?

There is no universal cutoff. Required ESS depends on the estimand, tail risk, desired interval precision, slice size, and cost of error. Predefine a decision-specific threshold through simulation or historical calibration and also inspect maximum weights and support.

Can ESS be used directly to calculate a confidence interval?

Not by default. Replacing `n` with ESS in an independent-sample formula ignores how weights were estimated, outcome heterogeneity, clustering, time dependence, and estimator structure. Use a variance or resampling method designed for the actual estimator and sampling unit.

Related Terms