What is Conformal Prediction?
Conformal Prediction is a model-agnostic statistical framework that converts point predictions or model scores into prediction sets or intervals with a declared finite-sample marginal coverage guarantee under exchangeability assumptions.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Freeze the score, split, and quantile convention
A conformal contract includes the base model and preprocessing, training and calibration split, Nonconformity Score, target miscoverage alpha, corrected empirical quantile, tie handling, output construction, and evaluation slices. Reusing training residuals for calibration breaks the ordinary split-conformal protocol. The Angelopoulos and Bates tutorial derives the finite-sample correction and covers regression intervals, classification sets, conditional diagnostics, and distribution shift.
Evaluate coverage together with efficiency
Coverage alone is trivial to satisfy with an interval spanning every value or a set containing every class. Report empirical coverage with uncertainty plus mean and tail interval width, prediction-set size, singleton and empty-set rates, latency, and downstream utility. Slice by class, language, difficulty, source, and time. Conditional coverage for every possible feature value is not supplied by the ordinary marginal guarantee, so large subgroup failures can coexist with acceptable aggregate coverage.
Monitor exchangeability and recalibrate after drift
Temporal dependence, feedback loops, active sampling, policy changes, class-prevalence shifts, and changed sensors can violate or weaken the exchangeability argument. Preserve labels in arrival order, test rolling coverage, and define refresh or fallback rules before release. Current MAPIE documentation separates prediction intervals and sets, risk control, conditional methods, and exchangeability testing; these extensions have different assumptions and must not be described as the basic split-conformal guarantee.
Key Characteristics
- Wraps many existing predictors without requiring a parametric outcome distribution
- Uses held-out calibration Nonconformity Scores and a corrected empirical quantile
- Produces prediction intervals for regression or label sets for classification
- Provides finite-sample marginal coverage under exchangeability
- Does not automatically provide individual, conditional, subgroup, or shifted-distribution coverage
- Must balance coverage with interval width, set size, latency, and downstream utility
Common Use Cases
- Adding coverage-controlled intervals to regression forecasts
- Returning plausible label sets for uncertain classification cases
- Quantifying uncertainty around document, image, or speech model outputs
- Supporting abstention or human review when sets are too large or intervals too wide
- Auditing coverage drift across time, language, source, or user-group slices
Example
Loading code...Frequently Asked Questions
What guarantee does Conformal Prediction provide?
Under the declared conformal procedure and exchangeability of calibration and future examples, the true outcome is contained in the prediction set or interval with at least the target marginal probability. It is not a per-input correctness guarantee.
Does distribution-free mean Conformal Prediction works under any drift?
No. Distribution-free means no parametric outcome distribution is required. The ordinary guarantee still relies on exchangeability or another explicitly justified assumption. Time dependence, feedback, policy changes, and population shift require separate methods and monitoring.
What is the difference between Conformal Prediction and probability calibration?
Probability calibration aligns numeric probabilities with observed frequencies. Conformal Prediction uses ranked Nonconformity Scores to construct sets or intervals with a coverage guarantee. The base probabilities need not be calibrated, and conformal output is not automatically a probability.
How should a conformal predictor be evaluated?
Measure empirical coverage with uncertainty and pair it with interval width or set size, singleton and empty-set rates, latency, and task utility. Repeat by class, language, source, difficulty, subgroup, and time so marginal coverage cannot hide local failures.
Can Conformal Prediction be used with LLMs?
It can wrap a clearly scored task to form answer sets, risk-controlled retrieval or abstention policies, but the score, outcome space, calibration data, and exchangeability assumptions must be explicit. It does not certify free-form truth, safety, or authorization by itself.