What is Path Patching?
Path Patching is a mechanistic interpretability intervention that changes the contribution carried from selected sender components to selected receiver components while keeping other modeled paths at a reference state, so a specific edge or path can be tested.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Represent the model as a computation graph
Choose nodes such as attention heads, MLPs, residual positions, or finer Query-Key-Value terms, and define directed edges as contributions between them. Select matched clean and corrupted inputs plus a signed scalar metric. Interpretability in the Wild introduced path patching while tracing the Indirect Object Identification circuit in GPT-2 Small.
Isolate a sender-to-receiver route
Cache both reference runs. Construct a patched computation in which the selected sender's contribution into the chosen receiver comes from the alternate run, while the sender's effects on non-target receivers and unrelated inputs to the receiver remain controlled. Multi-edge paths require ordered interventions and explicit rules for every branch; implementation shortcuts can silently test a different graph.
Validate paths against alternatives
Measure raw and normalized effects across held-out prompt pairs, multiple corruption methods, token positions, and seeds. Compare node patching, direct-logit attribution, random edges, and alternative paths. Activation-patching best-practice research shows that corruption and metric choices can materially change localization, so path results inherit those sensitivities plus extra interaction and off-distribution risks.
Key Characteristics
- Tests selected edges or routes rather than replacing every effect of a component
- Uses a declared computation graph with explicit sender and receiver nodes
- Requires matched clean and corrupted runs plus a signed behavior metric
- Controls non-target paths to isolate mediated information flow
- Can expose redundancy, negative paths, compensation, and route interactions
- Remains sensitive to graph granularity, corruption design, and patch order
Common Use Cases
- Testing whether one attention head influences another through a specific residual path
- Separating a sender's direct output effect from effects mediated by later components
- Tracing candidate information routes backward from an output metric
- Comparing competing circuit diagrams under matched interventions
- Validating edges retained by automated circuit-discovery methods
Example
Loading code...Frequently Asked Questions
How is Path Patching different from Activation Patching?
Activation Patching replaces a node activation and allows that change to propagate through every downstream route. Path Patching attempts to alter only the contribution from a selected sender to a selected receiver or path while controlling other routes, so it can test a narrower information-flow claim.
Why can a Path Patching restoration score exceed 100 percent?
Isolating one route can remove an opposing or compensating effect that remains present in the complete clean computation. Nonlinear interactions and denominator choice can also create overshoot. Always report clean, corrupted, and patched raw metrics alongside any normalized restoration value.
What must be fixed before running a Path Patching experiment?
Fix the model revision, token alignment, clean and corrupted distributions, graph decomposition, sender, receiver, patch direction, metric sign, normalization, aggregation, and random seeds. Changing any of these can change which path appears important and what the score means.
Can Path Patching identify a unique model circuit?
No. Parallel or redundant routes can compensate for one another, and the chosen graph granularity may split or merge real computations. Path Patching tests specified routes under specified interventions. Circuit claims also need faithfulness, completeness, minimality, alternative-path, and distribution checks.
When should Path Patching follow Direct Logit Attribution?
DLA can cheaply identify components that write directly toward an output direction. Path Patching can then test whether upstream components influence those writers through particular routes. This staged workflow is useful, but low DLA does not rule out a sender with a strong indirect effect.