What is Activation Patching?

Activation Patching is a causal intervention method that replaces selected internal activations in one model run with activations from a controlled comparison run and measures how the replacement changes a predefined output behavior.

Quick Facts

SpecificationOfficial Specification

How It Works

Construct a controlled pair of runs

Choose clean and corrupted inputs that isolate the variable under study while keeping syntax, length, and unrelated semantics as stable as possible. Cache internal activations from both runs, then replace one site or a defined group of sites. Sites may be residual-stream positions, attention outputs, MLP outputs, heads, neurons, or learned feature directions.

Measure restoration or disruption

A metric may be a target logit difference, probability, loss, ranking, or task score. Normalize recovery only when the clean-corrupt gap is meaningful and report the raw values as well. Causal tracing in model editing research combines corruption and hidden-state restoration to localize factual associations, illustrating the intervention pattern behind many patching studies.

Control for off-manifold and interaction effects

Patched activations can create states the model would not naturally produce, and correlated components can substitute for each other. Compare noising and denoising directions, alternative corruptions, resampling or mean baselines, multiple metrics, and negative-control positions. Single-site scans miss interactions; group or path patching can test combinations but expand the search space and still do not guarantee a complete circuit.

Key Characteristics

  • Transfers a selected internal activation between controlled model runs
  • Supports denoising tests that restore behavior and noising tests that disrupt it
  • Requires an explicit corruption, patch site, replacement source, and output metric
  • Can operate at layer, token, head, neuron, or learned-feature granularity
  • Provides causal localization under a specific intervention rather than a complete explanation
  • Is sensitive to off-manifold states, redundancy, interactions, and baseline choice

Common Use Cases

  1. Localizing where a model represents a subject, relation, or answer-relevant signal
  2. Testing which token positions and layers mediate a controlled behavior
  3. Comparing candidate attention heads, MLPs, or sparse-autoencoder features
  4. Validating a proposed mechanistic circuit with targeted interventions
  5. Prioritizing model components for deeper analysis or carefully scoped editing

Example

loading...
Loading code...

Frequently Asked Questions

What is the difference between denoising and noising Activation Patching?

Denoising inserts a clean activation into a corrupted run and asks how much target behavior returns. Noising inserts a corrupted activation into a clean run and asks how much behavior disappears. They are not interchangeable because redundancy, nonlinear interactions, and off-manifold states can make the two directions asymmetric.

How is Activation Patching different from ablation?

Ablation usually removes a component or replaces it with a generic baseline such as zero, mean, or resampled activation. Activation Patching uses a state from a specific comparison run, preserving more structured information and testing whether that run-specific state changes the selected behavior.

Does a high patching score identify the model's complete circuit?

No. It shows that replacing the selected state has a large effect under one experiment. Other paths may be redundant, multiple components may interact, and the patch may carry several variables at once. Circuit claims require compositional tests, controls, and coverage beyond a single heatmap.

How should the corrupted input be chosen?

Change the causal variable of interest while matching unrelated properties such as syntax, length, tokenization, and task difficulty as closely as possible. Use several corruption families and verify that the clean-corrupt metric gap reflects the intended behavior rather than a broad distribution shift.

What metrics are appropriate for Activation Patching?

Use a metric tied directly to the hypothesis, such as a target-versus-distractor logit difference, task loss, probability, or ranking. Report clean, corrupted, and patched values, define any normalization denominator, and check whether conclusions persist under reasonable alternative metrics.

Related Terms