What is Tuned Lens?
The Tuned Lens is an interpretability method that trains a separate affine translator for each model layer so an intermediate residual state can be decoded through the model's unembedding into a distribution that approximates the model's final token distribution.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Train one translator per layer
Collect residual states and the frozen model's final output distribution for a training corpus. For layer l, learn an affine map A_l h_l + b_l, then apply the checkpoint's final normalization and unembedding. The Tuned Lens objective distills the final model distribution, usually with cross-entropy or Kullback-Leibler divergence, rather than fitting human task labels.
Evaluate on held-out prompts and baselines
Keep translator training, hyperparameter selection, and evaluation examples disjoint. Report held-out divergence, entropy, calibration, target-token rank, seed variation, and performance by layer and prompt group. Compare with the raw Logit Lens, identity or mean baselines, and capacity-matched controls. A translator that memorizes corpus patterns can look accurate without faithfully revealing a particular activation.
Interpret translated trajectories cautiously
Improved prediction of the final distribution supports the claim that a learned affine map can decode that layer, not that the model itself applies the map. The translator may rotate features, recover information distributed across dimensions, or inject its own prior. Use the lens for hypothesis generation and diagnostics; verify claims about computation with activation patching, ablation, or other controlled interventions.
Key Characteristics
- Learns a distinct affine translator for each inspected model layer
- Keeps the source language model frozen while training translators
- Usually distills the model's final next-token distribution rather than task labels
- Can reduce intermediate-to-output basis mismatch relative to the raw Logit Lens
- Requires held-out data, capacity controls, calibration checks, and seed reporting
- Remains a learned observational probe and does not establish causal model use
Common Use Cases
- Tracing how a model's final token distribution becomes recoverable across layers
- Comparing intermediate prediction trajectories across prompts or checkpoints
- Diagnosing where uncertainty, alternatives, or factual candidates are resolved
- Generating candidate layers and token positions for causal intervention
- Auditing whether raw Logit Lens patterns persist after learned basis correction
Example
Loading code...Frequently Asked Questions
What problem does the Tuned Lens solve?
The final unembedding may be a poor direct decoder for early residual states because representations change basis across depth. The Tuned Lens learns a per-layer affine map that makes those states more predictive of the final distribution. It reduces a decoding mismatch, not necessarily a model-computation mismatch.
How is a Tuned Lens translator trained?
The language model is frozen. Intermediate states and the model's final token distributions are collected on a training corpus, and each layer's affine translator minimizes a distillation loss such as cross-entropy or KL divergence. Translator training and evaluation prompts must remain separate.
Can a Tuned Lens add information that is not in the activation?
An affine map cannot reconstruct arbitrary missing information, but its learned bias and rotation can encode corpus-level priors and emphasize distributed correlations. A high-capacity or poorly controlled probe may therefore predict well without matching the model's own computation. Capacity controls and held-out groups matter.
How should Tuned Lens quality be evaluated?
Report held-out divergence from the final distribution, entropy, calibration, target-token rank, sample count, seed variation, and results by layer and prompt group. Compare with the raw Logit Lens and simple or capacity-matched controls. Never select checkpoints or claims using the final test set.
Does a Tuned Lens trajectory explain why the model answered?
Not by itself. It shows what a trained translator can recover from each layer and can generate useful mechanistic hypotheses. To explain causal contribution, intervene on the implicated states or components, preserve appropriate controls, and test a predeclared behavior metric across a representative dataset.