What is Permutation Feature Importance (PFI)?

Permutation Feature Importance (PFI) is a model-inspection method that measures the change in a fitted model's evaluation score after one feature's values are permuted across a chosen dataset.

Quick Facts

SpecificationOfficial Specification

How It Works

Define the model-reliance experiment

Freeze the trained model and preprocessing pipeline, select representative labeled evaluation data, and declare a loss or score. Compute the baseline, permute one feature while preserving all other columns, and recompute the metric. research on model reliance formalizes permutation-based importance and emphasizes that reliance belongs to a particular prediction model rather than to the variable alone.

Repeat permutations and preserve the metric direction

Use multiple reproducible permutations and report the distribution, not only a rank. For losses where lower is better, importance is permuted loss minus baseline loss; for scores where higher is better, reverse the subtraction. Prefer held-out or cross-fitted evaluation when measuring generalization reliance, and report raw units so a value can be compared with the baseline performance.

Diagnose dependence and redundant information

Marginal shuffling breaks both the feature-target relation and dependencies with other inputs. A correlated substitute can keep performance high even when the inspected feature is used, while implausible shuffled rows can exaggerate degradation. Inspect correlation and model quality, permute meaningful groups, stratify by relevant cohorts, or use a validated conditional sampler when the intended question is unique information beyond other features.

Key Characteristics

  • Measures a fitted model's score change after disrupting one feature
  • Works with any model that can score a labeled evaluation dataset
  • Captures the inspected feature's main and interaction contributions to performance
  • Requires repeated permutations to quantify Monte Carlo variation
  • Changes with the metric, evaluation population, model, and permutation scheme
  • Can understate redundant features or evaluate unrealistic correlated combinations

Common Use Cases

  1. Ranking which inputs a deployed model relies on for held-out performance
  2. Comparing feature reliance across model versions, cohorts, or time windows
  3. Detecting features used for training fit but not for generalization
  4. Testing grouped importance for encoded variables or correlated feature families
  5. Triaging features for deeper PDP, ALE, SHAP, or domain review

Example

loading...
Loading code...

Frequently Asked Questions

What does Permutation Feature Importance measure?

It measures how much a fixed model's chosen evaluation metric changes when one feature is shuffled in a specific dataset. A larger degradation means the model relied more on information disrupted by that permutation. It does not measure a feature's universal value, causal effect, or importance to every other well-performing model.

Should PFI be computed on training or held-out data?

Use held-out or cross-fitted data when the question is which features support generalization. Training-set PFI instead measures reliance for fitting those seen observations and can make memorized noise appear useful. In either case, report the split, population, model version, metric, repeats, and uncertainty so the result is reproducible.

Why should permutations be repeated?

A single shuffle is one random reassignment and may produce an unusually mild or severe score change, especially with small datasets. Repeat the operation with controlled seeds or a documented permutation set, then report the mean and spread. Repetition quantifies shuffle variability; it does not correct distribution shift or feature dependence.

How do correlated features affect permutation importance?

A correlated feature can substitute for the shuffled feature, making either one appear weak even when the model uses both. Conversely, marginal shuffling can create implausible combinations and exaggerate the score loss. Grouped, subgroup, or conditional permutation can answer different questions, but the chosen dependency model must be validated and disclosed.

How is PFI different from SHAP?

PFI measures global performance degradation on labeled data after feature information is disrupted. SHAP allocates a selected prediction relative to a background value under a declared coalition-value convention and can operate locally. Their magnitudes and rankings need not agree because they use different estimands, outputs, data requirements, and dependency assumptions.

Related Terms