What is Classification Accuracy?

Classification Accuracy is the fraction of evaluated examples whose predicted hard label exactly matches the declared reference label.

Quick Facts

SpecificationOfficial Specification

How It Works

Compare with the correct no-skill baseline

For imbalanced single-label data, the majority-class predictor can achieve high Accuracy without finding any minority examples. Report that baseline, the complete Confusion Matrix, and per-class Recall. If error costs differ, a cost-weighted decision metric may be more relevant than assigning every correct label the same value.

Define what one correct example means

Ordinary single-label Accuracy gives each evaluated example one vote. Sample weights change that estimand. In multilabel classification, the current scikit-learn API uses subset Accuracy: every label for an example must match exactly. Token, span, document, conversation, and user-level Accuracy therefore answer different questions even when their formulas look identical.

Protect the evaluation boundary

Join predictions to labels by stable example ID, keep invalid outputs and abstentions visible, split data by the entity or time boundary that must generalize, and freeze the test set after model selection. Add confidence intervals or paired uncertainty when comparing releases. Accuracy on training data, a leaked holdout, or incomplete delayed labels is not evidence of production generalization.

Key Characteristics

  • Counts exact hard-label matches over the evaluated sample or sample weight
  • Uses all classes but assigns equal value to every correct prediction by default
  • Depends on the decision threshold for score-based binary classifiers
  • Can be dominated by a frequent majority class
  • Uses exact-set matching under common multilabel subset semantics
  • Does not measure ranking quality, probability calibration, or unequal error costs

Common Use Cases

  1. Checking a balanced single-label classifier against a simple baseline
  2. Tracking exact routing-label matches after the label contract is frozen
  3. Comparing deterministic classifiers at an identical operating threshold
  4. Measuring exact-set correctness for declared multilabel outputs
  5. Reporting a coarse summary alongside classwise errors and uncertainty

Example

loading...
Loading code...

Frequently Asked Questions

How is Classification Accuracy calculated?

Divide the number or total weight of exact hard-label matches by all evaluated examples or total weight. Publish the label set, sampling unit, weighting, threshold, invalid-output policy, and confidence interval so the denominator and meaning remain reproducible.

Why can high Accuracy be misleading?

When one class dominates, predicting that class for every example can look accurate while completely missing the minority class. Compare with the majority baseline and inspect the Confusion Matrix, per-class Recall, Precision, support, and decision costs.

What is multilabel subset Accuracy?

It counts an example as correct only when its complete predicted label set exactly matches the reference set. This is stricter than averaging correctness over individual labels, so reports must name the multilabel convention rather than saying only Accuracy.

Does Accuracy evaluate predicted probabilities?

No. Accuracy evaluates hard labels after an argmax or threshold. Two models can have identical labels and Accuracy while assigning very different probabilities. Use Log Loss or Brier Score for probability quality and ROC-AUC or PR analysis for ranking.

What should accompany Accuracy in a production report?

Include the baseline, Confusion Matrix, per-class metrics, sample support, uncertainty, important slices, threshold and abstention behavior, label maturity, drift checks, and the operational cost of each error type.

Related Terms

Related Articles