What is Expected Calibration Error?
Expected Calibration Error is a binned estimator that summarizes the weighted absolute difference between average model confidence and observed accuracy or event frequency across groups of predictions.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Fix the estimator identity before comparing releases
Record the event being calibrated, score used as confidence, population and time window, number and type of bins, boundary convention, weighting, empty-bin policy, and whether the calculation is top-label, classwise, marginal, or task-specific. The widely used formulation in On Calibration of Modern Neural Networks groups maximum class probabilities into equal-width bins, but that implementation is not a universal definition for every probabilistic output.
Inspect reliability instead of trusting one scalar
Binning creates bias and variance. Too few bins can hide local overconfidence; too many leave noisy or empty groups. Top-label ECE can also ignore probabilities assigned to non-winning classes and can conceal failures in rare classes or languages. Plot mean confidence against observed frequency with sample counts and intervals, then report signed gaps and critical slices. The AISTATS evaluation framework documents why calibration depends on the chosen notion and estimator.
Pair ECE with utility and proper probability scores
Use Brier Score or Log Loss to evaluate the full probability forecast, and report accuracy, ranking, abstention, and business loss separately. Compare a candidate with the same frozen labels, bins, model output definition, and bootstrap or repeated-sample uncertainty. Do not adopt a universal ECE threshold: acceptable error depends on sample support, decision cost, subgroup risk, and whether deployment data matches the evaluation distribution.
Key Characteristics
- Aggregates weighted absolute confidence-versus-frequency gaps across bins
- Depends on bin count, boundaries, weighting, and empty-bin handling
- Often evaluates only maximum-class confidence unless another variant is declared
- Can be low for an uninformative but frequency-matched predictor
- Is not a proper scoring rule and does not replace accuracy or business loss
- Needs reliability plots, sample counts, uncertainty, and slice analysis
Common Use Cases
- Summarizing classifier reliability under a frozen evaluation protocol
- Detecting overconfidence after model, prompt, or domain changes
- Comparing post-hoc calibrators on the same held-out predictions
- Auditing LLM answer-confidence or judge probabilities against adjudicated labels
- Monitoring high-confidence bins that feed automation or human-review routing
Example
Loading code...Frequently Asked Questions
How is Expected Calibration Error calculated?
Group predictions into declared confidence bins, compute each bin's mean confidence and observed accuracy, take the absolute gap, and weight it by the bin's share of samples. Publish bin boundaries, boundary rules, empty-bin handling, and the exact confidence definition.
Does a low ECE mean the model is accurate?
No. A predictor can always emit the base rate and be well calibrated while having weak discrimination. Pair ECE with task accuracy or loss, proper probability scores, ranking metrics, and decision outcomes.
Why does ECE change when the number of bins changes?
ECE approximates an underlying calibration relationship with grouped samples. Coarse bins average away local errors, while fine bins increase sampling noise and empty groups. Compare releases only with an identical, documented binning protocol.
What is the difference between ECE and Brier Score?
ECE is a bin-dependent summary of observed calibration gaps. Brier Score is a strictly proper score over each probability and outcome, combining calibration, resolution, and outcome uncertainty. Neither should be interpreted as the other.
What ECE threshold is acceptable for production?
There is no universal threshold. Set limits from representative sample support, decision cost, high-confidence errors, subgroup behavior, and a baseline under the same protocol. A small aggregate ECE cannot override a critical slice failure.