What is Partial Dependence Plot (PDP)?

Partial Dependence Plot (PDP) is a global model-inspection method that shows the average prediction of a fitted model as one feature or a small feature set is fixed to selected values while the remaining features follow a reference dataset.

Quick Facts

SpecificationOfficial Specification

How It Works

Specify the partial-dependence estimand

Choose the fitted model, exact scalar output, feature set, reference rows, sample weights, and grid before computing the curve. For a grid value, overwrite the selected feature in every reference row and average the resulting predictions. Friedman's gradient-boosting paper introduced this partial-dependence construction as a model-interpretation tool.

Read the average together with heterogeneity

A one-feature PDP shows an average response curve; a two-feature PDP can expose a joint surface. Because a PDP is the average of Individual Conditional Expectation curves over the same rows and grid, opposing subgroup effects can cancel into a flat line. Plot ICE curves, subgroup PDPs, and data density alongside the average, and state whether the output is a logit, probability, score, or regression value.

Test support, stability, and causal boundaries

When the varied feature is correlated with other inputs, replacement can produce combinations absent from the joint distribution. Restrict the grid to supported quantiles, inspect impossible rows, compare subgroups and Accumulated Local Effects, and repeat on relevant evaluation populations. The curve describes interventions on the model input; without an identified causal model, it does not show the real-world effect of changing that feature.

Key Characteristics

  • Summarizes a fitted model's average response to one feature or a feature pair
  • Marginalizes over a declared reference dataset rather than fitting a new model
  • Can reveal nonlinear shape, thresholds, saturation, and low-order interactions
  • Equals the average of ICE curves when both use the same rows and grid
  • Depends on target output, output scale, grid, weights, and reference population
  • Can evaluate unsupported feature combinations and does not establish causality

Common Use Cases

  1. Checking whether a production model learned an expected monotonic response
  2. Comparing average feature-response curves across model versions or cohorts
  3. Inspecting a two-feature interaction before defining domain-specific tests
  4. Finding regions where sparse data make a model response difficult to trust
  5. Pairing an average curve with ICE or ALE to diagnose hidden heterogeneity

Example

loading...
Loading code...

Frequently Asked Questions

What exactly does a Partial Dependence Plot estimate?

It estimates the fitted model's average selected output after setting the feature of interest to each grid value across a declared reference dataset. The average is over the remaining observed feature values. It is conditional on the model, preprocessing, output scale, rows, weights, and grid, so changing any of them can change the curve.

What is the difference between a PDP and an ICE plot?

An ICE plot draws one response curve for each reference row while the selected feature varies. A PDP averages those curves at every grid value. ICE can reveal subgroups or interactions whose positive and negative responses cancel in the average, while PDP offers a compact global summary that is easier to compare.

Why are correlated features a problem for partial dependence?

Replacing one feature while leaving correlated features unchanged can create combinations that the model never saw and that may be impossible in the domain. Predictions on those rows still enter the average. Inspect joint support, limit the grid, stratify the population, and compare ALE or conditional approaches before trusting the shape.

Should a classification PDP use probabilities or logits?

Either can be valid, but they answer different questions. Probabilities are familiar but the nonlinear link function can compress effects near zero and one; logits preserve additive changes on the score scale. Declare the class, target contrast, output unit, and any calibration step, and do not compare curves computed on different scales.

Does a PDP show the causal effect of changing a feature?

Not by itself. It shows how the fitted model responds when an input is replaced according to the PDP procedure. The model may encode confounding, proxies, selection bias, or impossible interventions. A real-world causal claim requires a causal estimand, an identification strategy, valid assumptions, and data supporting that intervention.

Related Terms