What is Balanced Accuracy?
Balanced Accuracy is the unweighted mean of Recall calculated for every declared class, giving each class equal influence regardless of its sample support.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Understand the class-balanced weighting identity
Balanced Accuracy is equivalent to ordinary weighted Accuracy when each example receives a weight inversely proportional to the support of its true class. This is an evaluation identity, not a recommendation to use those weights during training. The current scikit-learn API defines the score as mean per-class Recall and offers a separate chance-adjusted option.
Keep absent classes and support visible
A declared class with no test examples has undefined Recall. Silently dropping it changes the metric's class set, while forcing zero may answer a different question. Freeze the label inventory before evaluation, report support and per-class Recall, and ensure the test set covers every class that the product promises to handle. Rare-class uncertainty remains high even when each class has equal weight.
Separate class balance from decision costs
Brodersen and colleagues motivate Balanced Accuracy for biased classifiers on imbalanced data and derive a posterior distribution under stated assumptions. Equal class weighting is still not the same as equal business harm. Compare thresholds on explicit false-positive and false-negative costs, and add Precision, calibration, workload, and slice evidence.
Key Characteristics
- Averages Recall equally across every declared class
- Equals the mean of Sensitivity and Specificity in binary classification
- Reduces majority-class dominance in the evaluation aggregate
- Is equivalent to Accuracy under inverse true-class-frequency sample weights
- Requires an explicit policy for classes absent from the evaluation sample
- Does not rebalance training data or encode business-specific error costs
Common Use Cases
- Comparing classifiers when class frequencies differ substantially
- Giving rare and common routing labels equal evaluation influence
- Detecting majority-class predictors that achieve misleading Accuracy
- Monitoring class-balanced performance after prevalence shifts
- Reporting a global score alongside per-class Recall and uncertainty
Example
Loading code...Frequently Asked Questions
How is Balanced Accuracy calculated?
Compute Recall separately for every declared class and take their unweighted arithmetic mean. In binary classification this becomes the mean of Sensitivity and Specificity. Publish the class set, each support count, weighting, and missing-class policy.
How does Balanced Accuracy differ from ordinary Accuracy?
Ordinary Accuracy gives each example equal influence, so common classes dominate. Balanced Accuracy gives each true class equal influence through mean class Recall. The two are equal when class supports are equal or classwise Recall happens to be identical.
Does Balanced Accuracy fix an imbalanced training dataset?
No. It changes how evaluation results are aggregated; it does not resample data, alter the loss, improve labels, or repair a biased classifier. Training interventions and evaluation metrics are separate choices that require independent validation.
What happens when a class is absent from the test set?
That class has undefined Recall, so the declared Balanced Accuracy cannot be estimated as written. Do not silently remove the class. Report the missing support, improve the evaluation sample, or apply a prespecified policy and label the changed estimand.
Is Balanced Accuracy enough for production approval?
No. It hides which classes fail and treats classes as equally important. Add the Confusion Matrix, per-class Recall and Precision, support and uncertainty, calibration or ranking metrics where relevant, and actual error costs at the deployed threshold.