What is Accumulated Local Effects (ALE)?
Accumulated Local Effects (ALE) is a global model-inspection method that averages prediction differences within data-supported feature intervals, accumulates those local effects, and centers the result to describe a fitted model's feature response.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Estimate local differences inside supported intervals
Partition the selected feature, commonly by quantiles. Within each interval, take rows observed there, replace only that feature with the lower and upper boundaries, and average the prediction difference. The original ALE paper designed this conditional local calculation to avoid extrapolating all reference rows to every grid value.
Accumulate and center the effects
Sum interval effects from a baseline boundary to each point, then subtract the sample-weighted mean so the first-order ALE curve averages to zero. A value above zero means the model prediction is higher than its average contribution at that feature value on the selected output scale. It is not the prediction itself. Second-order ALE additionally removes both main effects and isolates an interaction surface.
Stress-test bins and interpretation
Quantile bins balance row counts but can span very different numeric widths; too few bins oversmooth, while too many create noisy or empty intervals. Report bin boundaries and counts, repeat across resolutions and samples, inspect missing cells in two-dimensional ALE, and compare PDP or ICE where appropriate. ALE reduces correlation-driven extrapolation but does not make the fitted association causal.
Key Characteristics
- Averages upper-minus-lower prediction changes within observed feature intervals
- Uses the conditional distribution of remaining features inside each interval
- Accumulates local effects and centers the resulting first-order curve at zero
- Reduces off-support extrapolation caused by correlated features in PDP
- Supports first-order feature effects and centered second-order interactions
- Depends on bin resolution, sample support, output scale, and reference population
Common Use Cases
- Inspecting model response when a feature is strongly correlated with other inputs
- Comparing feature-effect shape with a PDP to detect extrapolation artifacts
- Finding supported ranges where the model response changes direction
- Examining pure two-feature interactions after removing main effects
- Monitoring feature-response drift across model versions or evaluation cohorts
Example
Loading code...Frequently Asked Questions
How are Accumulated Local Effects computed?
Split the feature range into supported intervals. For rows observed in each interval, evaluate the fitted model after replacing that feature with the interval's upper and lower boundaries, then average the prediction differences. Accumulate these local averages across intervals and subtract a weighted constant so the final first-order curve has mean zero.
Why can ALE be preferable to a Partial Dependence Plot?
PDP replaces a feature with every grid value for all rows, which can create implausible combinations when features are correlated. ALE computes changes only for rows observed near each interval, so it stays closer to the joint data support. It reduces this extrapolation problem, but sparse intervals and conditional sampling still need inspection.
What does a positive ALE value mean?
After centering, a positive first-order ALE value means the selected feature value contributes a higher fitted response than the feature's average contribution over the reference population, on the declared output scale. It is not the full prediction, a probability unless that scale was selected, or a feature attribution for one individual.
How many ALE bins should be used?
There is no universal count. Use enough intervals to preserve relevant shape while retaining adequate observations per interval. Quantile boundaries often balance counts, but may span uneven widths. Report boundaries and counts, repeat the calculation at several resolutions, and treat unstable tails or empty two-dimensional cells as unsupported.
Does ALE solve every problem caused by correlated features?
No. ALE reduces PDP-style extrapolation by conditioning local differences on observed intervals, but it can still be noisy under weak support, sensitive to bins, and difficult to interpret when features are nearly deterministic functions of one another. It describes the fitted model and does not identify a real-world causal effect.