What is F1 Score?
F1 Score is the harmonic mean of Precision and Recall for a declared positive class at a specific decision threshold, reaching a high value only when both metrics are high.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Use F-beta when error priorities are unequal
The F-beta family is (1+beta^2)PR/(beta^2 P+R). Values above one weight Recall more; values below one weight Precision more. van Rijsbergen's information-retrieval treatment derives the composite effectiveness family from an explicit user preference between Recall and Precision. Beta encodes relative emphasis, not a complete business cost model; operational thresholds still need explicit utility and constraints.
Declare positive class, threshold, and zero policy
F1 is attached to hard predictions at one threshold. If there are no true positives, false positives, or false negatives, its denominator is zero and the score is undefined. Libraries may warn, return zero, return one, or omit the class depending on configuration. The current scikit-learn API makes label selection, averaging, beta, support, and zero-division behavior explicit.
Do not collapse multiclass behavior prematurely
Binary F1 evaluates one declared positive class. Macro-F1 averages classwise F1 equally, Weighted-F1 uses true-class support, and Micro-F1 aggregates TP, FP, and FN before scoring. In single-label multiclass classification, Micro-F1 equals Accuracy, so it can still hide weak minority classes. Inspect per-class scores and the confusion matrix before choosing an aggregate.
Key Characteristics
- Combines Precision and Recall through their harmonic mean
- Uses true positives, false positives, and false negatives but ignores true negatives
- Depends on the positive-class definition and one decision threshold
- Extends to F-beta when Precision and Recall receive different emphasis
- Requires an explicit policy when its denominator is zero
- Has distinct Binary, Macro, Micro, Weighted, and Samples aggregation semantics
Common Use Cases
- Summarizing positive-class detection when false positives and misses both matter
- Comparing text or entity classifiers on a frozen label set
- Selecting a validation threshold under an explicitly symmetric error objective
- Reporting Macro-F1 when every declared class should influence model selection
- Detecting regressions while retaining Precision, Recall, support, and slice details
Example
Loading code...Frequently Asked Questions
How is the F1 Score calculated?
F1 is `2 * Precision * Recall / (Precision + Recall)`, equivalently `2TP/(2TP+FP+FN)`. The two forms agree when their denominators are defined. Publish the positive class, threshold, counts, and zero-division policy with the result.
Why does F1 use the harmonic mean?
The harmonic mean is pulled toward the smaller component, so high Precision cannot fully compensate for very low Recall, or vice versa. This gives a compact balance score, but it does not prove that the two error types have equal business cost.
What is the difference between F1 and F-beta?
F1 sets beta to one and gives Precision and Recall symmetric emphasis. F-beta above one emphasizes Recall, while beta below one emphasizes Precision. The chosen beta must come from the decision objective and should be reported rather than hidden in a generic score.
Why can a high F1 Score still be misleading?
F1 ignores true negatives and compresses two error dimensions into one number. A classifier may perform badly on the negative class, a critical subgroup, calibration, or workload while retaining a high F1. Inspect confusion counts, per-class metrics, slices, and costs.
What is the difference between Macro-F1 and Micro-F1?
Macro-F1 computes F1 per class and averages classes equally. Micro-F1 first aggregates TP, FP, and FN across classes. In single-label multiclass classification, Micro-F1 equals Accuracy, while Macro-F1 gives rare classes the same influence as common classes.