What is Linear Probing?
Linear Probing is an analysis method that trains a linear classifier or regressor on frozen internal representations to measure whether specified information can be decoded under a documented data, feature, optimization, and evaluation protocol.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Define the representation and readout
Specify the model checkpoint, component, layer, token position, pooling rule, preprocessing, target label, and linear model. Freeze the source model and fit only the readout. A linear classifier tests whether a hyperplane separates classes; a linear regressor tests whether a weighted sum predicts a continuous target. Feature standardization, intercepts, regularization, and class weights are part of the protocol.
Build splits, baselines, and controls
Split by the unit that can leak, such as document, template, speaker, subject, or sequence, rather than randomly splitting correlated rows. Compare against majority, metadata-only, raw-input, random-feature, and untrained-model baselines where appropriate. Hewitt and Liang propose control tasks and selectivity to distinguish useful linguistic structure from a probe's ability to memorize arbitrary labels.
Keep conclusions within the evidence
Held-out performance establishes decodability under one distribution, not necessity, sufficiency, localization, or downstream use by the model. Check confidence intervals, seed variation, regularization sweeps, sample-efficiency curves, and out-of-distribution groups. To claim causal use, follow the probe with controlled interventions such as activation patching, ablation, or steering and measure task behavior.
Key Characteristics
- Uses frozen model activations as features for a constrained linear readout
- Requires explicit choices of layer, token position, pooling, labels, and regularization
- Depends on leakage-resistant train, validation, and test splits
- Needs simple baselines, control tasks, and uncertainty estimates
- Measures linear decodability rather than unique representation or model use
- Supports comparison across layers, checkpoints, datasets, and training stages
Common Use Cases
- Comparing where task-relevant information becomes linearly decodable across layers
- Monitoring representation changes during training or fine-tuning
- Testing whether demographic, stylistic, or domain attributes are exposed in embeddings
- Screening candidate layers before more expensive causal interventions
- Auditing whether a reported probe result survives grouped splits and control tasks
Example
Loading code...Frequently Asked Questions
What does a successful Linear Probe prove?
It proves that the chosen target can be decoded by the specified linear readout from the selected activations on the evaluated distribution. It does not prove that the information is represented only there, that the model uses it for its output, or that the representation causes the behavior.
Why must probe splits be grouped?
Activations from the same document, template, person, or sequence can be highly correlated. A random row split may place near-duplicates on both sides and inflate test accuracy. Grouping by the true leakage unit gives a more credible estimate of generalization to unseen entities or contexts.
What are control tasks and probe selectivity?
A control task assigns arbitrary labels while preserving nuisance structure, testing whether the probe can memorize rather than recover the intended property. Selectivity compares performance on the real task with the control task. It is useful evidence about probe capacity, not a complete validity certificate.
Should a Linear Probe always use the same regularization?
No. Tune regularization only on validation data and report the search range, selected value, sample size, and seed variation. Layer comparisons should use equivalent tuning budgets. A result that appears only at one extreme setting may reflect optimization or capacity rather than a stable representation.
How can a probe result support a causal claim?
Use the probe to identify a candidate representation, then intervene on that representation while preserving suitable controls. Activation patching, ablation, or steering can test whether changing it changes the relevant behavior. Even then, define the metric and alternative pathways before claiming necessity or sufficiency.