What is Direct Logit Attribution?
Direct Logit Attribution is a mechanistic interpretability method that projects individual residual-stream component outputs onto a chosen unembedding or logit-difference direction to estimate each component's direct additive contribution to that output metric.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Decompose the residual stream
At a fixed token position, express the final residual state as the sum of embeddings, attention-head outputs, MLP outputs, and relevant bias terms. The residual-stream view in A Mathematical Framework for Transformer Circuits makes this additive decomposition explicit. Decide whether to group by block, attention layer, head, neuron, or another component before comparing scores.
Project onto a declared output direction
For a target token, use its unembedding vector; for a contrastive task, use the difference between target and distractor unembedding vectors. Take each component's inner product with that direction after applying a documented approximation to the final normalization. TransformerLens exposes residual decomposition and logit-attribution helpers, but tensor positions and normalization conventions still require validation.
Separate direct evidence from causal influence
A positive score means a component directly raises the chosen logit metric under the decomposition; a negative score lowers it. A near-zero direct score does not imply irrelevance because the component may alter later queries, keys, values, gates, or MLP features. Compare prompts, positions, baselines, and seeds, then use path patching, activation patching, or ablation to test downstream causal effects.
Key Characteristics
- Uses the additive residual stream to assign signed direct output contributions
- Projects components onto a token or target-minus-distractor unembedding direction
- Can decompose results by layer, attention head, MLP, neuron, or embedding
- Requires an explicit treatment of final normalization, bias, and token position
- Omits indirect effects mediated through later components
- Produces attribution evidence rather than a complete or uniquely causal circuit
Common Use Cases
- Finding attention heads or MLP blocks that directly support a target token
- Comparing positive and negative contributions to a contrastive logit difference
- Tracing how residual-stream writes accumulate across model depth
- Selecting candidate senders for path patching or exact activation interventions
- Checking whether an apparent component effect is stable across prompts and positions
Example
Loading code...Frequently Asked Questions
What does a Direct Logit Attribution score mean?
It is the signed projection of one residual-stream component onto the selected output direction under a stated normalization convention. Positive values support the target relative to the contrast token; negative values oppose it. Magnitude is metric-specific and should not be compared across incompatible prompts or scales.
How is Direct Logit Attribution different from the Logit Lens?
The Logit Lens decodes the accumulated residual state at successive layers into full vocabulary distributions. DLA decomposes one selected final logit or logit difference into direct contributions from individual writes such as heads and MLPs. One tracks states; the other attributes an output direction.
How should LayerNorm be handled in DLA?
LayerNorm is not globally linear because its scale depends on the combined residual state. Analysts commonly apply the final run's normalization scale to each component or use a library helper that documents the approximation. The contributions must be checked against reconstruction of the selected final metric.
Can a component with near-zero DLA still matter?
Yes. DLA captures only the component's direct projection into the output direction. A component can strongly influence later attention patterns, MLP gates, or features while writing little direct evidence itself. Path patching, activation patching, and ablation are needed to test those mediated effects.
Does a large DLA score prove a component is causal?
No. The score is an algebraic attribution under a chosen decomposition, and correlated components can write similar directions. Establish causal importance by intervening on the component or path, checking matched controls and alternative metrics, and confirming that the effect generalizes beyond discovery prompts.