What is Equalized Odds?
Equalized Odds is a group fairness criterion requiring a prediction to be statistically independent of protected-group membership conditional on the observed outcome, which for binary classification means equal true-positive and false-positive rates across groups.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Compare errors within each observed outcome
Compute TPR=TP/(TP+FN) among labeled positives and FPR=FP/(FP+TN) among labeled negatives for every group. Report both gaps because equal sensitivity can coexist with unequal false alarms. Hardt, Price, and Srebro introduced Equalized Odds and a post-processing construction that can use group membership to transform a score or classifier.
Treat labels as measurements, not unquestioned truth
Conditional error rates inherit every flaw in the label. If only selected applicants later receive an outcome, if policing intensity changes recorded incidents, or if access affects diagnosis, the evaluated label process differs by group. Audit missingness, selection, adjudication agreement, time horizon, and intervention effects before interpreting a smaller gap as fairer treatment.
Evaluate thresholds, randomization, and tradeoffs
Different group thresholds or randomized post-processing can reduce measured gaps but may be restricted, operationally unstable, or inconsistent with the product's policy and legal context. Equalized Odds can also conflict with Predictive Parity when observed base rates differ and predictions are imperfect. Compare utility and harms, confidence intervals, calibration, subgroup support, and downstream actions rather than optimizing one maximum gap.
Key Characteristics
- Conditions the fairness comparison on the observed outcome label
- Requires equal true-positive and false-positive rates across groups
- Includes Equal Opportunity as a common one-class relaxation
- Depends critically on label validity, availability, timing, and consistency
- Can be approximated through training constraints, thresholds, or post-processing
- May conflict with calibration or predictive-parity objectives
Common Use Cases
- Comparing qualified-candidate recall and unqualified-candidate advancement
- Auditing false denial and false approval patterns in credit decisions
- Evaluating clinical alert sensitivity and false-alarm burden by group
- Monitoring threshold changes with both favorable and adverse error costs
- Testing whether a post-processor transfers beyond its validation population
Example
Loading code...Frequently Asked Questions
How is Equalized Odds measured?
Within each group, calculate true-positive rate among observed positives and false-positive rate among observed negatives. Compare both rates across groups and report counts and uncertainty. The metric is undefined for a group-label stratum with no observations, so software should not silently replace that result with zero.
What is the difference between Equalized Odds and Equal Opportunity?
Equalized Odds constrains both true-positive and false-positive rates. Equal Opportunity usually constrains only the true-positive rate for a designated beneficial outcome. The relaxation can fit contexts where missed benefits dominate, but it intentionally leaves the other error rate unconstrained.
Can Equalized Odds be trusted when labels are biased?
Only with strong caution. The metric treats observed labels as the conditioning reference. If labels reflect selective observation, unequal access, prior enforcement, or inconsistent adjudication, equal conditional rates may preserve those distortions. Label-process audits are part of the fairness assessment.
Why can Equalized Odds conflict with Predictive Parity?
With unequal observed outcome prevalences and imperfect prediction, equal true-positive and false-positive rates generally imply different positive predictive values. Equalizing positive predictive value instead generally leaves some conditional error rates unequal. Teams must document which error or decision property matters.
Does using different thresholds by group solve Equalized Odds?
It can reduce measured disparities on a validation set, but it does not settle whether group-specific treatment is lawful, acceptable, stable, or beneficial. Thresholds can drift, labels can be incomplete, and downstream burdens can change. Review policy, legal, uncertainty, and operational evidence together.