What is Counterfactual Explanation?
Counterfactual Explanation is a local explanation that describes one or more constrained changes to an instance's features that would make a fitted model produce a specified alternative output.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Define the desired output and action set
Declare the instance, model version, exact output target or tolerance, features a person or operator can change, immutable attributes, monotonic constraints, valid categories, time horizon, and action costs. Wachter, Mittelstadt, and Russell framed counterfactual explanations as small changes that obtain a desired model outcome without exposing the model's internal logic.
Optimize several quality criteria
A useful candidate must achieve the target, remain close under a meaningful distance, change few features, stay plausible under the joint data distribution, and satisfy feasibility constraints. Because several distinct actions can reach the same output, return a diverse set rather than one arbitrary optimum. DiCE research formalizes the tradeoff between proximity and diversity while allowing contextual constraints.
Validate recourse instead of promising outcomes
Recheck candidates through the exact production preprocessing and model, test nearby perturbations and model versions, estimate data support, and review whether actions are available at comparable cost across protected groups. A counterfactual is a statement about the fitted decision function. Claiming that the recommended action will cause the desired real outcome requires a structural causal model and intervention assumptions beyond ordinary counterfactual explanation.
Key Characteristics
- Targets a declared alternative class, score, probability, or output interval
- Searches for nearby feature changes under explicit feasibility constraints
- Balances validity, proximity, sparsity, plausibility, cost, and diversity
- Can be generated with model-specific gradients or model-agnostic search
- Must preserve immutable attributes and valid dependencies between actions
- Explains a model decision but does not guarantee causal or stable recourse
Common Use Cases
- Showing an applicant which controllable changes would alter a model decision
- Testing whether a decision boundary offers realistic options for affected users
- Discovering implausible shortcuts or brittle regions near a prediction
- Comparing recourse cost and availability across cohorts or model versions
- Generating bounded what-if cases for human review and contestability
Example
Loading code...Frequently Asked Questions
What makes a counterfactual explanation valid?
At minimum, the candidate must produce the declared target output when passed through the exact production preprocessing and model. Operational validity also requires allowed values, immutable-feature constraints, feasible transitions, acceptable cost, joint-distribution support, and a stated time horizon. Passing the model threshold alone is necessary but not sufficient.
How is a counterfactual explanation different from feature attribution?
Attribution methods such as SHAP assign contribution values to features for an observed prediction under a defined background convention. A counterfactual searches for an alternative input that changes the prediction. Attribution explains contribution under one configuration; counterfactuals answer a target-seeking what-if question and require action and feasibility constraints.
Is a counterfactual explanation the same as an adversarial example?
Both can cross a model decision boundary, but their objectives differ. An adversarial example often seeks an imperceptible perturbation that exposes vulnerability, even if the change is not actionable. A counterfactual explanation should be understandable, feasible, sparse, plausible, and relevant to the user's goal, although poor generators can still return adversarial artifacts.
Why should a system return multiple counterfactuals?
Decision boundaries usually admit several solutions, and the mathematically nearest one may be unavailable or undesirable for a particular user. A diverse set exposes tradeoffs among cost, time, sparsity, and controllable features. Diversity must be constrained by validity and plausibility; visually different but infeasible candidates do not provide meaningful choice.
Does following a counterfactual guarantee the real outcome?
No. The explanation guarantees only that the configured model predicts the target for the proposed input, subject to numerical tolerance. Real actions can affect other variables, the model may be wrong or retrained, and the input-output relationship may be confounded. A causal guarantee needs an identified structural model, intervention semantics, and outcome validation.