What is Integrated Gradients?

Integrated Gradients is a feature-attribution method that multiplies each input-to-baseline difference by the integral of a selected model output's gradient along a path from the baseline to that input.

Quick Facts

SpecificationOfficial Specification

How It Works

Define the target, baseline, and path

Choose an exact scalar output such as a class logit, probability, loss, or target-minus-reference score. Then choose a baseline that represents absence or a meaningful reference in the modeled input space. The original Integrated Gradients paper uses a straight path from baseline to input, but different credible baselines can answer different contrastive questions and produce different attributions.

Approximate the path integral

Interpolate points between baseline and input, compute the target gradient at each point, aggregate with a declared Riemann or quadrature rule, and multiply by the feature-wise displacement. Increase steps until the attribution vector and convergence delta stabilize. Current Captum documentation exposes multiple numerical rules, step counts, internal batching, and a completeness delta.

Validate attribution beyond completeness

Check that attribution sum approximates target(input) minus target(baseline), then vary baselines, step counts, paths, seeds, preprocessing, and feature grouping. Use deletion, insertion, or domain-valid perturbations to test whether highly ranked features affect the declared output. A small completeness error only verifies the numerical decomposition; correlated features, unrealistic baselines, and off-manifold paths can still make the interpretation misleading.

Key Characteristics

  • Attributes a selected scalar output relative to a declared baseline input
  • Integrates gradients along a path instead of using only the endpoint gradient
  • Usually applies a straight-line interpolation from baseline to observed input
  • Satisfies Sensitivity and Implementation Invariance under the method definition
  • Offers a completeness check against the baseline-to-input output difference
  • Remains sensitive to baseline semantics, target units, path, and approximation error

Common Use Cases

  1. Attributing image-class logits to pixels relative to one or more reference images
  2. Scoring token-embedding dimensions or grouped tokens for a selected text output
  3. Diagnosing gradient saturation that hides important input changes
  4. Comparing explanation stability across baselines and numerical step counts
  5. Checking whether suspected features affect a differentiable model output

Example

loading...
Loading code...

Frequently Asked Questions

How are Integrated Gradients different from ordinary input gradients?

An ordinary saliency gradient measures local sensitivity only at the observed input and can be near zero in a saturated region. IG averages gradients along a baseline-to-input path and scales by the input difference. That captures accumulated change, but it introduces baseline and path assumptions that endpoint gradients do not have.

How should an Integrated Gradients baseline be chosen?

Choose a reference whose absence or comparison meaning is defensible for the application and preprocessing. Zero, black, padding, or an average sample can each be misleading in some domains. Compare several plausible baselines or a baseline distribution, disclose the choice, and avoid treating attribution outside the data manifold as neutral.

What does the completeness property guarantee?

Under exact integration, the feature attributions sum to the selected output at the input minus that output at the baseline. A numerical implementation reports a residual or convergence delta. A small delta confirms the decomposition is numerically consistent; it does not show that the baseline, target, or feature semantics are appropriate.

Can Integrated Gradients explain text models?

Yes, but discrete tokens do not form a natural continuous path. Implementations usually interpolate embeddings relative to a padding, zero, or reference embedding and then aggregate dimensions by token. The tokenizer, embedding layer, special tokens, baseline sequence, target position, and aggregation rule must all be recorded.

Do Integrated Gradients establish causal feature importance?

No. IG attributes a model-output difference along a mathematical path and can expose sensitivity under that setup. It does not identify the real-world cause of an outcome, and correlated features can redistribute meaning. Use domain-valid interventions, counterfactual design, held-out tests, and causal assumptions for stronger claims.

Related Terms

Related Articles