What is Permutation Feature Importance (PFI)?
Permutation Feature Importance (PFI) is a model-inspection method that measures the change in a fitted model's evaluation score after one feature's values are permuted across a chosen dataset.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Define the model-reliance experiment
Freeze the trained model and preprocessing pipeline, select representative labeled evaluation data, and declare a loss or score. Compute the baseline, permute one feature while preserving all other columns, and recompute the metric. research on model reliance formalizes permutation-based importance and emphasizes that reliance belongs to a particular prediction model rather than to the variable alone.
Repeat permutations and preserve the metric direction
Use multiple reproducible permutations and report the distribution, not only a rank. For losses where lower is better, importance is permuted loss minus baseline loss; for scores where higher is better, reverse the subtraction. Prefer held-out or cross-fitted evaluation when measuring generalization reliance, and report raw units so a value can be compared with the baseline performance.
Diagnose dependence and redundant information
Marginal shuffling breaks both the feature-target relation and dependencies with other inputs. A correlated substitute can keep performance high even when the inspected feature is used, while implausible shuffled rows can exaggerate degradation. Inspect correlation and model quality, permute meaningful groups, stratify by relevant cohorts, or use a validated conditional sampler when the intended question is unique information beyond other features.
Key Characteristics
- Measures a fitted model's score change after disrupting one feature
- Works with any model that can score a labeled evaluation dataset
- Captures the inspected feature's main and interaction contributions to performance
- Requires repeated permutations to quantify Monte Carlo variation
- Changes with the metric, evaluation population, model, and permutation scheme
- Can understate redundant features or evaluate unrealistic correlated combinations
Common Use Cases
- Ranking which inputs a deployed model relies on for held-out performance
- Comparing feature reliance across model versions, cohorts, or time windows
- Detecting features used for training fit but not for generalization
- Testing grouped importance for encoded variables or correlated feature families
- Triaging features for deeper PDP, ALE, SHAP, or domain review
Example
Loading code...Frequently Asked Questions
What does Permutation Feature Importance measure?
It measures how much a fixed model's chosen evaluation metric changes when one feature is shuffled in a specific dataset. A larger degradation means the model relied more on information disrupted by that permutation. It does not measure a feature's universal value, causal effect, or importance to every other well-performing model.
Should PFI be computed on training or held-out data?
Use held-out or cross-fitted data when the question is which features support generalization. Training-set PFI instead measures reliance for fitting those seen observations and can make memorized noise appear useful. In either case, report the split, population, model version, metric, repeats, and uncertainty so the result is reproducible.
Why should permutations be repeated?
A single shuffle is one random reassignment and may produce an unusually mild or severe score change, especially with small datasets. Repeat the operation with controlled seeds or a documented permutation set, then report the mean and spread. Repetition quantifies shuffle variability; it does not correct distribution shift or feature dependence.
How do correlated features affect permutation importance?
A correlated feature can substitute for the shuffled feature, making either one appear weak even when the model uses both. Conversely, marginal shuffling can create implausible combinations and exaggerate the score loss. Grouped, subgroup, or conditional permutation can answer different questions, but the chosen dependency model must be validated and disclosed.
How is PFI different from SHAP?
PFI measures global performance degradation on labeled data after feature information is disrupted. SHAP allocates a selected prediction relative to a background value under a declared coalition-value convention and can operate locally. Their magnitudes and rankings need not agree because they use different estimands, outputs, data requirements, and dependency assumptions.