What is Confusion Matrix?
Confusion Matrix is a table whose rows and columns represent actual and predicted classes, so each evaluated example contributes to exactly one cell and exposes which classes a classifier gets right or confuses.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Read raw counts before normalized views
Raw counts reveal sample support and operational workload. Row normalization answers how predictions are distributed within each actual class and supports Recall-style analysis; column normalization answers how actual labels are distributed within each predicted class and supports Precision-style analysis. Global normalization shows population shares. The current scikit-learn API exposes all three choices, but normalization must not replace the underlying denominators.
Bind every matrix to a threshold and label contract
A scoring classifier produces a different matrix at each decision threshold. Lowering the positive threshold usually moves cases from the predicted-negative column to the predicted-positive column, changing both benefits and errors. In multiclass evaluation, inspect the full K x K matrix and per-class one-versus-rest views; a diagonal aggregate can hide a systematic confusion between two costly classes.
Audit the evidence behind actual labels
The matrix treats the supplied label as truth, but labels can be delayed, selectively observed, inconsistent, or changed by the intervention triggered by a prediction. Join predictions to labels by stable example ID, preserve abstentions and invalid outputs, fix an outcome-maturity window, and report uncertainty by deployment slice. Fawcett's ROC analysis shows how the binary table underlies threshold-specific rates and ROC operating points.
Key Characteristics
- Cross-tabulates actual classes against predicted classes
- Preserves separate counts for each correct and incorrect class pairing
- Changes when a score threshold, abstention rule, or label mapping changes
- Supports raw, row-normalized, column-normalized, and globally normalized views
- Extends from a 2 x 2 binary table to a K x K multiclass table
- Requires explicit axis orientation, label order, support, and evaluation population
Common Use Cases
- Finding which support intents a routing model confuses
- Separating false alarms from missed fraud or safety events
- Comparing model releases at the same operating threshold
- Auditing class-specific behavior in imbalanced multiclass data
- Checking errors by language, device, region, or other deployment slice
Example
Loading code...Frequently Asked Questions
How do you read a Confusion Matrix?
First identify which axis is actual and which is predicted, then confirm class order and the positive class. In a binary matrix, inspect TP, FP, TN, and FN as raw counts before calculating rates. Diagonal cells are correct only under the displayed ordering; never infer orientation from position alone.
Should a Confusion Matrix show counts or percentages?
Publish raw counts because they preserve support and workload, then add a clearly labeled normalized view for the question at hand. Row percentages support Recall analysis, column percentages support Precision analysis, and whole-matrix percentages describe population share. One view cannot replace the others.
Why does a Confusion Matrix change with the threshold?
A threshold converts scores into class labels. Moving it changes which examples are predicted positive or negative, so examples move between cells. Record the model, score definition, threshold, positive class, and abstention policy with every matrix.
How is a multiclass Confusion Matrix interpreted?
Rows and columns represent the same declared class order. Each off-diagonal cell identifies one actual-to-predicted confusion. Inspect raw support, row-normalized recall, column-normalized precision, and high-cost class pairs instead of relying only on the diagonal total.
Can a Confusion Matrix prove that a classifier is ready for production?
No. It describes labeled cases under one protocol. Production evidence also needs representative splits, uncertainty, calibration or ranking checks where relevant, latency, abstentions, invalid outputs, slice analysis, delayed-label handling, and business consequences at the chosen operating point.