What is Log Loss?

Log Loss is the mean negative logarithm of the probability a classifier assigns to each observed label, so confident wrong predictions receive a much larger penalty than uncertain ones.

Quick Facts

SpecificationOfficial Specification

How It Works

Compute from valid probabilities or stable logits

Probability vectors must be finite, non-negative, aligned to the declared class order, and sum to one within tolerance. Evaluation libraries commonly clip probabilities away from zero and one for numerical safety; publish that epsilon because it caps the largest penalty. During training, use a numerically stable logits-based implementation such as LogSumExp instead of applying Softmax and then taking a raw logarithm.

Use a proper score without calling it calibration

The logarithmic score is strictly proper: in expectation, the true predictive distribution minimizes its negative value. Gneiting and Raftery develop the general proper-scoring-rule framework. A lower aggregate Log Loss does not by itself prove calibration on every probability range or subgroup; add reliability analysis and task outcomes.

Preserve weighting, label, and split contracts

The current scikit-learn API uses natural logarithms, supports binary and multiclass probabilities, and clips to machine precision. Class weights, sample weights, label smoothing, missing outcomes, and renormalization all change the measured objective. Evaluate on untouched data and compare models under identical class order, preprocessing, weights, and label horizon.

Key Characteristics

  • Scores the probability assigned to the observed class
  • Penalizes confident incorrect probabilities logarithmically without a finite theoretical cap
  • Is strictly proper under a correctly specified outcome space and expectation
  • Supports binary and multiclass probability distributions
  • Depends on logarithm base, aggregation, weighting, clipping, and class order
  • Does not select a decision threshold or prove slice-level calibration

Common Use Cases

  1. Comparing probabilistic classifiers that share the same label space
  2. Training logistic-regression and neural classification models
  3. Detecting releases that become more confidently wrong
  4. Evaluating reward-model or judge probabilities against adjudicated outcomes
  5. Auditing probability quality alongside calibration and threshold metrics

Example

loading...
Loading code...

Frequently Asked Questions

How is binary Log Loss calculated?

For each label `y` and positive probability `p`, calculate `-[y log(p)+(1-y)log(1-p)]`, then average using the declared sample weights. State the logarithm base, clipping epsilon, positive class, label window, and aggregation.

Why does Log Loss heavily penalize confident mistakes?

The negative logarithm grows without bound as the probability assigned to the observed class approaches zero. A wrong probability of 0.001 is therefore much worse than 0.4, even when both become the same incorrect hard label at a 0.5 threshold.

Is Log Loss the same as Cross-Entropy Loss?

For categorical outcomes represented by one-hot labels and predicted class probabilities, their standard empirical formulas coincide. Cross-entropy is a broader information-theoretic quantity, so a report should still state the target encoding, reduction, weights, logarithm base, and logits or probability input.

Does lower Log Loss prove better calibration?

No. Log Loss is a proper score affected by both probability reliability and the model's ability to separate outcomes. A lower aggregate can coexist with a weak probability bin or subgroup. Add reliability diagrams, calibration summaries, discrimination, and decision metrics.

Why do implementations clip probabilities?

Exact zero probability on an observed class gives infinite loss and can cause numerical failure. Clipping creates a finite computational value but changes the maximum penalty. Record epsilon and clipped counts, and prefer stable logits-based formulas during training.

Related Terms

Related Articles