What is Confidence Calibration?
Confidence Calibration is the property and process by which a model's stated probabilities are made consistent with observed outcome frequencies under a declared prediction, label, population, and evaluation protocol.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Measure probability meaning before changing scores
Use reliability diagrams to compare mean predicted probability with observed frequency, and report sample counts for every bin. Add a proper scoring rule such as Brier Score or Log Loss, discrimination metrics, and task-specific errors. Expected Calibration Error is a useful summary but depends on bin boundaries, sample size, and whether it evaluates top-label, classwise, or marginal calibration. The AISTATS calibration framework explains why different calibration definitions and estimators answer different questions.
Fit calibration without leaking evaluation labels
Post-hoc methods map frozen model scores to probabilities. Temperature Scaling divides multiclass logits by one learned positive temperature; Platt-style Sigmoid Calibration and Isotonic Regression make different shape and data assumptions. Fit the calibrator on data disjoint from model training and final testing, then freeze the base model, calibrator, label mapping, and preprocessing together. Current scikit-learn guidance uses held-out or cross-validated predictions because fitting on training predictions biases the mapping toward excessive confidence.
Treat deployment drift as a new calibration problem
Calibration is distribution-dependent. Changes in prevalence, language, device, prompt, model, retrieval policy, label definition, or reviewer population can invalidate an old mapping even when aggregate accuracy looks stable. Monitor reliability and decision cost by material slice, preserve an uncalibrated baseline, and recalibrate only with representative labels. A calibrated score can support thresholds or expected-cost decisions, but authorization, safety rules, and high-impact review remain deterministic external controls.
Key Characteristics
- Aligns probability values with empirical outcome frequencies under a declared protocol
- Is distinct from accuracy, ranking quality, uncertainty estimation, and factual correctness
- Requires held-out calibration data that represents the intended decision population
- Can use temperature scaling, sigmoid calibration, isotonic regression, or task-specific mappings
- Must be evaluated with reliability views, proper scores, sample support, and deployment slices
- Can fail after model, prompt, label, prevalence, or distribution changes
Common Use Cases
- Mapping classifier probabilities to decision thresholds with explicit error costs
- Calibrating reward-model or LLM-judge pairwise probabilities against human labels
- Routing uncertain document, image, or speech predictions to human review
- Comparing model releases without confusing probability quality with accuracy
- Monitoring overconfidence after domain, language, prompt, or policy changes
Example
Loading code...Frequently Asked Questions
What does it mean for a model to be calibrated?
Under a declared population and event definition, predictions assigned probability p should be correct or occur about p of the time. This is a group-frequency statement, not proof that any individual prediction has a known chance of being correct.
Is a calibrated model necessarily accurate?
No. A model that always predicts the base rate can be calibrated yet provide little discrimination. Report accuracy or task loss, ranking quality, calibration, and decision outcomes separately rather than using one metric as a substitute for all of them.
How do temperature scaling and isotonic regression differ?
Temperature scaling applies one positive scale to logits and preserves their class ordering. Isotonic regression learns a flexible monotone step function and can fit more shapes, but it needs more representative calibration data and can introduce probability ties.
Can an LLM's verbal confidence be treated as a calibrated probability?
Not without evaluation. A verbal percentage, token likelihood, self-rating, or judge score has a distinct generation process. Define the target event, collect representative outcomes, measure reliability and proper scores, and repeat after model or prompt changes.
When should a calibration mapping be retrained?
Re-evaluate it whenever the model, prompt, preprocessing, labels, population, language mix, prevalence, or decision policy changes, and when monitored reliability drifts. Use fresh representative labels and keep final test data out of calibrator fitting.