What is Brier Score?

Brier Score is a strictly proper scoring rule that averages the squared distance between a predicted probability distribution and the observed outcome encoded as a binary value or one-hot class vector.

Quick Facts

SpecificationOfficial Specification

How It Works

Interpret a score relative to an explicit baseline

A raw Brier value has no universal good threshold. Compare it with a frozen baseline such as the empirical base rate, the current production model, or another approved forecast on identical cases. Brier Skill Score is commonly expressed as 1 - candidate_score / reference_score; positive values indicate improvement over that reference, zero means parity, and negative values mean worse performance. The reference population and label window are part of the result.

Do not describe Brier Score as calibration alone

The score combines probability reliability with the forecast's ability to separate outcomes and the inherent outcome uncertainty. Current scikit-learn calibration guidance warns that a lower aggregate Brier loss can come from better discrimination even when calibration worsens. Use the Murphy decomposition where appropriate, plus Reliability Diagrams or ECE, discrimination metrics, and task errors before diagnosing a release.

Preserve probability and label contracts

Validate finite probabilities, class ordering, normalization, sample weights, event horizon, and delayed-label completeness. For LLM or Judge confidence, define how free-form answers become outcomes and how confidence is elicited; token likelihood and verbalized percentages are not interchangeable. Report the model, prompt, calibrator, dataset, language, and slice identity, and recompute after distribution or policy changes. The current scikit-learn API documents binary, multiclass, label-order, and scaling behavior explicitly.

Key Characteristics

  • Scores the full probability forecast against an observed outcome
  • Is strictly proper when probabilities and outcome space are correctly specified
  • Penalizes confident errors quadratically
  • Combines reliability, resolution, and outcome uncertainty in aggregate form
  • Has binary and multiclass scaling conventions that must be declared
  • Needs a reference score, task metrics, slices, and label-completeness checks

Common Use Cases

  1. Evaluating binary risk, fraud, failure, or success probabilities
  2. Scoring multiclass probability vectors instead of only top-one labels
  3. Comparing calibrated and uncalibrated model releases on frozen cases
  4. Auditing reward-model, LLM-judge, or verbalized answer confidence
  5. Computing skill relative to a base-rate or production reference forecast

Example

loading...
Loading code...

Frequently Asked Questions

How is the Brier Score calculated?

For a binary event, average the squared difference between each predicted positive-event probability and its zero-or-one outcome. For multiclass prediction, sum squared differences across the declared class vector for each case, then average.

What is a good Brier Score?

There is no universal cutoff. Compare the score with a base-rate, current-production, or other approved reference forecast on identical cases and scaling. Also inspect class balance, slices, uncertainty, and the cost of specific errors.

Is Brier Score a calibration metric?

It is a proper probability score influenced by calibration, but its aggregate also reflects resolution and outcome uncertainty. Use Reliability Diagrams, ECE or another declared calibration estimator when the question is specifically whether probabilities match observed frequencies.

Why do Brier Score ranges differ between tools?

Binary scoring often uses `(p-y)^2` and ranges from 0 to 1. The original multiclass form sums squared errors over classes and ranges from 0 to 2; some implementations scale selected cases by one half. Always publish the exact formula and option.

Can Brier Score evaluate LLM confidence?

Yes when each answer maps to a clearly adjudicated event and the confidence is a declared probability for that event. Preserve prompt, model, elicitation method, answer normalizer, label policy, and slices; do not substitute token likelihood without validation.

Related Terms

Related Articles