What is Qini Coefficient?

Qini Coefficient is an area-based metric that summarizes how far a model's cumulative incremental-outcome curve rises above the random-targeting baseline when units are ordered by predicted treatment effect.

Quick Facts

SpecificationOfficial Specification

How It Works

The curve adjusts treated responses by a control counterfactual

For binary outcomes, one common prefix gain is R_t(k)-R_c(k)N_t(k)/N_c(k), where R and N are responder and sample counts by arm among the top k ranked units. Plot that gain against targeted population share. The pylift documentation illustrates Qini curves, random baselines, practical maxima, and normalized variants.

Area above random measures ranking, not calibration

A model earns positive area when high-ranked prefixes contain more incremental response than random ordering. The same Qini score can arise from effect predictions with different numerical calibration, and a good ranking can still recommend economically harmful treatment when cost is ignored. Pair Qini with grouped-effect calibration, policy value at operating thresholds, treatment rate, and net benefit.

Use honest randomized or causally identified evaluation data

Computing the curve on training data rewards overfitting, while unadjusted observational treatment groups confound targeting quality with historical assignment. Use an untouched randomized holdout or valid cross-fitted causal scores, preserve known propensities and sample weights, quantify uncertainty with an appropriate bootstrap or influence method, and inspect curve stability across practical budget ranges.

Key Characteristics

  • Evaluates the ordering induced by predicted treatment effects
  • Builds a cumulative incremental-outcome curve over ranked prefixes
  • Subtracts a random-targeting reference through an area calculation
  • Depends on the declared control adjustment and normalization convention
  • Does not by itself measure effect calibration or net policy value
  • Requires honest causal evaluation data and uncertainty analysis

Common Use Cases

  1. Comparing uplift rankings on a randomized marketing holdout
  2. Selecting a contact threshold under a fixed campaign budget
  3. Checking whether a CATE model concentrates incremental conversions early
  4. Monitoring uplift-ranking degradation across cohorts or time
  5. Complementing policy-value and calibration diagnostics

Example

loading...
Loading code...

Frequently Asked Questions

How is the Qini Coefficient calculated?

Rank units by predicted uplift, estimate cumulative incremental outcome at successive population shares using treatment and control observations, integrate the resulting curve, and subtract the area under random targeting. Exact adjustment, interpolation, weighting, and normalization conventions must be reported.

What is the difference between Qini and AUUC?

Both summarize uplift curves, but naming and baselines vary across libraries. Qini commonly denotes area above a random-targeting line, while AUUC may denote total area under an uplift curve or an adjusted area. Inspect the implemented formula before comparing values or selecting models.

Can Qini Coefficients be compared across datasets?

Raw Qini usually scales with sample size, outcome prevalence, treatment ratio, and attainable gain, so cross-dataset comparisons are unsafe. A declared normalization can help, but changes in population, treatment cost, propensities, and evaluation design still require separate interpretation.

Does a positive Qini Coefficient prove a causal model is valid?

No. It supports useful ranking only under the evaluation design's causal assumptions. Training-set evaluation, hidden confounding, interference, poor overlap, or incorrect propensity weighting can create misleading curves. Statistical uncertainty and an honest holdout are essential.

Is Qini enough to choose a treatment policy?

No. Qini summarizes ranking over the whole curve, whereas deployment uses a particular threshold with costs, capacity, harm, and uncertainty. Evaluate policy value and net benefit at realistic treatment rates, inspect calibration and subgroups, and compare with simple operational baselines.

Related Terms