What is Attribution Patching?

Attribution Patching is a first-order approximation to activation patching that estimates how replacing a corrupted-run activation with its clean-run value would change a scalar metric, using the metric gradient at the corrupted activation dotted with the clean-minus-corrupted activation difference.

Quick Facts

SpecificationOfficial Specification

How It Works

Define a clean-corrupt-metric contract

Choose matched clean and corrupted inputs that differ in the feature under study while controlling other cues. Define a signed scalar metric before inspecting results, such as a target-minus-distractor logit difference. Cache comparable activations at the same component, layer, token position, and hook convention. The estimated patch direction changes if clean and corrupted states or metric sign are reversed.

Compute first-order patch estimates

Backpropagate the metric through the corrupted computation to obtain its gradient at each cached activation, then take the inner product with the clean-minus-corrupted activation difference. Attribution Patching makes this screening far cheaper than exact one-site-at-a-time patching. Aggregate only after preserving component and token detail needed by the hypothesis.

Verify finalists with exact interventions

Rank candidates across prompts, seeds, and corruption families, then rerun exact activation patches for the shortlisted sites. Compare signed effect, rank stability, and residual error. Large steps violate the local approximation; opposing feature effects can cancel in a dot product, and near-zero gradients can hide important nonlinear pathways. Methods such as AtP* address some known failure modes but do not remove the need for exact validation.

Key Characteristics

  • Approximates a clean-to-corrupt activation replacement with a first-order Taylor term
  • Uses gradients from a scalar metric evaluated around the corrupted run
  • Can score many layers, heads, neurons, tokens, or residual sites in one backward pass
  • Depends on matched clean and corrupted examples plus a signed metric contract
  • Can fail under saturation, nonlinearity, cancellation, and large activation differences
  • Serves as a screening method whose leading candidates require exact patch verification

Common Use Cases

  1. Ranking candidate attention heads or MLP sites before exact activation patching
  2. Scanning token positions and layers in large mechanistic interpretability studies
  3. Comparing candidate circuit locations across prompt and corruption families
  4. Estimating whether a clean activation could restore a corrupted behavior metric
  5. Auditing approximation quality by comparing predicted and exact patch effects

Example

loading...
Loading code...

Frequently Asked Questions

How is Attribution Patching different from Activation Patching?

Activation Patching actually replaces an internal state and reruns the remaining forward computation, producing an intervention effect for that patch. Attribution Patching estimates the same change with a local gradient dot activation difference, making broad scans cheaper but less reliable.

Which metric should Attribution Patching use?

Use a predeclared signed scalar tied to the behavior, such as target-minus-distractor logit difference, loss change, or calibrated task score. Avoid post hoc metrics selected because they make a component look important. Record direction, scale, aggregation, token position, and clean-corrupt ordering.

When does the first-order approximation fail?

It can fail when the clean-corrupt displacement is large, the computation is strongly nonlinear, gradients saturate near the corrupted point, or multiple effects cancel in the dot product. Normalization and interactions between components can also make one-site estimates differ from joint interventions.

How should Attribution Patching candidates be validated?

Repeat the screen across prompts, seeds, and corruption constructions, then perform exact patches on high-ranked candidates and a sample of low-ranked candidates. Compare effect sign, magnitude, rank stability, and false-negative rate. Use held-out examples rather than validating only on discovery prompts.

Is Attribution Patching the same as gradient descent?

No. It uses backpropagation to calculate a metric gradient, but it does not optimize model parameters by repeated updates. The gradient is multiplied by a measured clean-corrupt activation difference to approximate an intervention. Gradient descent is an optimization algorithm that changes parameters.

Related Terms