What is Logit Lens?
The Logit Lens is a training-free interpretability method that applies a language model's final normalization and unembedding to intermediate residual-stream states, producing vocabulary logits or probabilities that can be inspected across layers.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Project intermediate states into vocabulary space
Capture the residual stream at a documented hook point, such as before or after a block, and at a specified token position. Apply the same final LayerNorm or RMSNorm and unembedding matrix used by the checkpoint's output head. The original Logit Lens description uses this direct projection without fitting a separate translator.
Compare trajectories rather than isolated tokens
Inspect top-token rank, target-token logit, entropy, and divergence from the final distribution across layers. Keep tokenization, normalization, logit scale, and temperature fixed. Compare correct and contrastive prompts, multiple positions, random or untrained baselines, and repeated examples. A single appealing decoded token is weak evidence because vocabulary projections expose many correlated alternatives.
Account for basis mismatch and causal limits
Later layers may rotate, refine, or cancel directions before the final output, so directly unembedding an early residual state can mix genuine signal with basis mismatch. The Tuned Lens learns per-layer translators to reduce this mismatch, but introduces probe-training confounds. Neither lens establishes necessity or sufficiency; use activation patching, ablation, or another controlled intervention to test causal dependence.
Key Characteristics
- Reuses the checkpoint's final normalization and unembedding without training a probe
- Maps selected intermediate residual states to vocabulary logits or probabilities
- Requires exact hook-point, token-position, normalization, and tokenizer conventions
- Supports layerwise trajectories of token rank, entropy, logits, or divergence
- Can be distorted by intermediate-to-output basis mismatch, especially in early layers
- Provides observational evidence and does not by itself prove causal model use
Common Use Cases
- Inspecting how candidate next-token predictions evolve through model depth
- Comparing residual-stream trajectories for correct and contrastive prompts
- Finding layers where a target token becomes visible or is later suppressed
- Screening candidate components for activation patching or ablation experiments
- Checking whether claims from a single example persist across a controlled dataset
Example
Loading code...Frequently Asked Questions
What exactly is projected by the Logit Lens?
It projects an intermediate residual-stream vector at a chosen token position. The vector is passed through the checkpoint's final normalization and unembedding to produce vocabulary logits. Results depend on whether the hook is before or after a block and on the exact normalization convention.
Does the Logit Lens require training?
No separate probe is trained. It reuses the model's final output transformation, which makes it fast and avoids learned-probe capacity. That convenience does not make the projection unbiased: intermediate states may not yet be expressed in the basis expected by the final unembedding.
Why can early Logit Lens predictions be misleading?
Early residual states can contain features that later layers rotate, combine, suppress, or cancel. Applying the final decoder too soon may therefore expose basis mismatch or correlated token directions rather than a stable prediction. Dataset-level trajectories are more credible than one striking token.
How is the Logit Lens different from the Tuned Lens?
The Logit Lens directly applies the final output interface at every layer. The Tuned Lens first learns a separate affine translator for each layer, usually by matching the final model distribution. This can improve calibration but adds training data, optimization choices, and probe-induced information.
Can the Logit Lens prove that a layer causes an answer?
No. It shows that a token-related direction is visible under a particular projection. To test causal use, intervene on the candidate activation while controlling the rest of the computation, then measure a predeclared behavioral metric. Patching, ablation, or steering can provide stronger evidence.