What is LIME?
LIME (Local Interpretable Model-agnostic Explanations) is a post-hoc method that explains one black-box prediction by fitting a simple interpretable surrogate to model outputs on proximity-weighted perturbations around that instance.
Quick Facts
| Full Name | Local Interpretable Model-agnostic Explanations |
|---|---|
| Specification | Official Specification |
How It Works
Define an interpretable local neighborhood
Select the instance, target output, interpretable features, perturbation process, distance, and kernel. Text variants may toggle words, image variants may toggle superpixels, and tabular variants need a distribution that respects types and dependencies. The original LIME paper defines the explanation as a locally faithful interpretable model with a complexity penalty.
Query the black box and fit a sparse surrogate
Generate perturbed samples, obtain black-box scores rather than only hard labels, assign proximity weights, and fit a weighted linear model, short tree, or another declared interpretable family. Feature selection and regularization trade readability against local fidelity. The coefficients explain the fitted neighborhood in the interpretable representation; they are not global parameters of the original model.
Measure fidelity, stability, and plausibility
Report weighted holdout error around the target, repeat seeds, sweep kernel widths and sample counts, and compare alternative perturbation generators. Check whether samples remain plausible and whether the surrogate preserves the black box near relevant decision boundaries. Independent LIME guidance highlights how kernel width, distance, sampling, and instability can change the explanation.
Key Characteristics
- Explains one prediction with a simple model fitted in a weighted neighborhood
- Requires only prediction queries rather than model gradients or internal states
- Uses an interpretable representation that may differ from original model features
- Trades surrogate complexity against measured local fidelity
- Depends on perturbation sampling, distance, kernel width, and random seed
- Can generate plausible-looking coefficients from unrealistic local samples
Common Use Cases
- Explaining individual tabular predictions with a sparse local linear model
- Highlighting words that influence one text-classification score
- Selecting image superpixels associated with a particular prediction
- Comparing local behavior before and after a model release
- Auditing explanation instability across seeds and neighborhood definitions
Example
Loading code...Frequently Asked Questions
What does model-agnostic mean in LIME?
LIME needs a callable prediction interface but not gradients, weights, or architecture-specific internals. It can therefore wrap many model families. Model-agnostic does not remove assumptions: the interpretable representation, perturbation distribution, distance, kernel, target output, and surrogate family still define the explanation.
Is a LIME explanation local or global?
A standard LIME explanation is local to one instance and to the synthetic neighborhood defined by its kernel and sampling process. Its coefficients should not be read as global model behavior. Representative-instance selection can summarize several local explanations, but that still does not turn one surrogate into a global model.
Why can repeated LIME runs produce different explanations?
LIME commonly samples perturbations and may perform sparse feature selection, so random seed, sample count, correlated predictors, kernel width, and an unstable local fit can change selected features and weights. Repeat runs, report dispersion, test nearby inputs, and reject explanations whose fidelity or feature set is unstable.
How should LIME neighborhood quality be checked?
Verify feature types and constraints, inspect whether perturbations are plausible, hold out some weighted samples to measure local fidelity, sweep kernel width, and compare with real nearby observations when available. For text or images, confirm that token removal or superpixel masking does not create artifacts that dominate the black-box response.
Do LIME coefficients reveal causal effects?
No. They describe a surrogate fitted to predictions in a generated neighborhood. The coefficients depend on sampling, weighting, representation, and correlated features, and may include combinations that cannot occur in reality. Causal interpretation requires explicit interventions, identification assumptions, and domain-valid data beyond LIME.