What is Counterfactual Fairness?
Counterfactual Fairness is a causal fairness criterion requiring a predictor's distribution for the same individual to remain unchanged across counterfactual worlds that differ in a protected attribute while preserving the individual's inferred background conditions.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Declare the causal model and permissible pathways
Represent how the protected attribute, background factors, observed features, decisions, and outcomes are generated. Decide whether direct and mediated paths are prohibited, allowed, or unresolved; this is a normative and domain judgment, not a discovery produced by the fairness formula. The Counterfactual Fairness paper defines fairness through counterfactual invariance under a causal model.
Use abduction, action, and prediction
First infer a distribution over latent background conditions from the observed person. Next replace the structural assignment for the protected attribute with an alternative value. Then recompute descendants and the predictor under the modified model. Compare prediction distributions, not only point values, when background variables remain uncertain. Reusing the inferred background state is what makes the comparison concern the same individual.
Audit assumptions before trusting invariance
Multiple causal models can fit the same observed data yet produce different counterfactuals. Test alternative graphs, functions, latent-confounding assumptions, measurement models, and permissible-path policies; report sensitivity rather than one definitive number. Also retain group outcome and process audits because a counterfactually invariant predictor can still be inaccurate, inaccessible, poorly calibrated, or embedded in an unfair institution.
Key Characteristics
- Compares predictions for the same unit across protected-attribute interventions
- Requires a Structural Causal Model rather than association alone
- Preserves inferred exogenous background conditions across counterfactual worlds
- Recomputes causal descendants instead of editing one table cell in isolation
- Depends on normative decisions about permissible and impermissible pathways
- Needs sensitivity analysis because observational data rarely identify one unique model
Common Use Cases
- Auditing whether a risk score relies on protected-attribute causal pathways
- Comparing individual predictions under alternative modeled social conditions
- Testing proxy-removal claims that group-rate metrics cannot resolve
- Reviewing which descendant variables are permissible inputs to a decision
- Documenting causal assumptions behind a high-impact model fairness claim
Example
Loading code...Frequently Asked Questions
How is Counterfactual Fairness tested?
Specify an SCM, infer background conditions for an observed person, intervene on the protected attribute, recompute descendants, and compare the predictor's distribution across worlds. Repeat under plausible alternative models and report uncertainty because the required counterfactual quantities are generally not directly observed.
Why is changing only the protected column incorrect?
Protected attributes can causally affect observed descendants. Flipping one cell while freezing education, access, measurements, or other descendants may violate the model's equations and create an incoherent person. A causal intervention changes the selected mechanism and propagates effects through declared descendants.
Can a counterfactually fair model use descendants of a protected attribute?
Sometimes, but only under an explicit causal and normative treatment of pathways. A descendant may carry both permissible and impermissible effects, and observational data may not separate them. Simply including or excluding every descendant is not a universal solution.
Is Counterfactual Fairness the same as a Counterfactual Explanation?
No. Counterfactual Fairness compares a predictor across protected-attribute interventions for the same modeled individual. A Counterfactual Explanation searches for feasible input changes that achieve a target model output. The latter can be useful even without claiming protected-attribute invariance.
Can observational data prove Counterfactual Fairness?
Usually not by themselves. Different structural functions or latent-confounding assumptions can reproduce the same observed distribution and disagree on counterfactuals. Domain evidence, experiments where possible, model criticism, and sensitivity analysis are required, and conclusions remain conditional on the declared SCM.