What is Causal Forest?
Causal Forest is a forest-based estimator that builds adaptive covariate neighborhoods for estimating Conditional Average Treatment Effects, with splitting and inference designed around treatment-effect variation rather than outcome prediction alone.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Trees learn an adaptive treatment-effect neighborhood
Each tree recursively partitions covariate space with a criterion related to heterogeneity in the target quantity. Across many subsampled trees, records that repeatedly share leaves receive larger forest weights for a query point. The final CATE comes from a local estimating equation over those weighted neighbors. The GRF algorithm reference documents this adaptive-nearest-neighbor view and treatment-specific splitting.
Honesty and orthogonalization address different biases
Honesty uses different observations to choose splits and populate leaves, reducing adaptive estimation bias at the cost of data efficiency. Modern implementations can residualize outcomes and treatment with out-of-bag nuisance predictions, following an R-learner-style orthogonalization that reduces sensitivity to first-stage errors. Neither technique fixes unmeasured Confounding, weak Overlap, or a poorly defined intervention.
Validate the forest as an effect model, not a classifier
Outcome RMSE and feature importance do not establish accurate CATE. Use out-of-bag or held-out grouped calibration, best linear projection, heterogeneity or rank-weighted tests, policy-value evaluation, variance estimates, and stability across seeds and tuning choices. Check treatment counts inside local neighborhoods, unsupported predictions, cluster or time dependence, and whether the discovered heterogeneity transfers.
Key Characteristics
- Targets Conditional Average Treatment Effects rather than outcomes alone
- Uses an ensemble to form adaptive local covariate neighborhoods
- Can separate split selection from effect estimation through honesty
- Often orthogonalizes treatment and outcome with nuisance predictions
- Supports out-of-bag effects and statistical inference under conditions
- Still depends on causal identification, overlap, and transportability
Common Use Cases
- Exploring high-dimensional treatment-effect heterogeneity in experiments
- Estimating CATE without prespecifying every interaction
- Ranking units for a capacity-constrained treatment policy
- Testing whether baseline features predict incremental product impact
- Generating interpretable subgroup hypotheses for a new confirmatory trial
Example
Loading code...Frequently Asked Questions
How is a Causal Forest different from a Random Forest?
A Random Forest typically minimizes outcome-prediction error. A Causal Forest targets a local treatment contrast, uses treatment-aware splitting and estimating equations, and may orthogonalize nuisance components. Feeding treatment into an ordinary predictor and subtracting two predictions is a different estimator.
What does honesty mean in a Causal Forest?
Honesty means observations used to choose a tree's partition are separated from observations used to estimate effects in its leaves. This reduces adaptive bias and supports inference under stated conditions, but it reduces the effective data available for each task and can hurt very small samples.
Does a Causal Forest estimate Individual Treatment Effects?
It estimates CATE at a feature vector, an average over a learned local population. The two realized Potential Outcomes for one unit remain unobserved. Calling the output an exact individual effect confuses personalized prediction with identifiable person-level truth.
Can Causal Forests be used with observational data?
Yes, if treatment, outcome, timing, adjustment set, conditional Exchangeability, Positivity and dependence assumptions are credible. Propensity and outcome nuisance models can improve estimation, but the forest cannot diagnose or remove hidden confounding by itself.
How should Causal Forest predictions be evaluated?
Use out-of-bag or held-out grouped effects, calibration and heterogeneity tests, rank-weighted effects or policy value, confidence intervals, overlap diagnostics, and stability across seeds and tuning. Ordinary outcome accuracy and attractive subgroup plots are insufficient.