What is Superposition?
Superposition is a neural representation strategy in which a model encodes more features than the available activation dimensions by assigning features to non-orthogonal directions, accepting interference when those features are read or combined.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Capacity pressure favors non-orthogonal features
An activation vector with d dimensions has at most d mutually orthogonal directions, yet a model may need to represent far more context-dependent features. When features are sparse, the model can encode them along nearly orthogonal directions and tolerate occasional collisions. The Toy Models of Superposition study demonstrates how feature importance, sparsity, and interference can produce this geometry in controlled networks.
Interference separates features from neurons
A feature is a behaviorally meaningful direction or pattern, while a neuron is one coordinate in the implementation basis. A single feature may span many neurons, and one neuron may participate in many features. Reading a feature with a dot product also picks up cross-terms from non-orthogonal directions, so neuron labels and raw activation magnitudes can conceal the represented structure.
Diagnose geometry without overclaiming
Useful evidence includes feature sparsity, pairwise direction overlap, reconstruction quality, intervention effects, and stability across datasets or checkpoints. Sparse autoencoders can propose a larger feature dictionary, but their latents depend on the training objective and dictionary size. Superposition is one explanation for distributed representations, not a guarantee that every mixed activation has a clean feature decomposition.
Key Characteristics
- Encodes more candidate features than the number of activation dimensions
- Uses non-orthogonal feature directions and therefore permits interference
- Becomes more favorable when features are sparse or rarely co-active
- Separates implementation coordinates such as neurons from represented features
- Can produce polysemantic units without making polysemanticity a complete diagnostic
- Requires geometric, causal, and cross-context evidence rather than visual inspection alone
Common Use Cases
- Explaining why individual neurons respond to several unrelated concepts
- Analyzing capacity and interference in compact neural representations
- Motivating overcomplete sparse dictionaries for activation analysis
- Comparing representation geometry across layers, models, or training regimes
- Designing interventions that target feature directions instead of single neurons
Example
Loading code...Frequently Asked Questions
Is neural-network Superposition the same as quantum superposition?
No. In neural-network interpretability, the term describes encoding several features in non-orthogonal directions of an activation space. It is a geometric and optimization phenomenon in an ordinary numerical model, not a quantum state or a claim about quantum computation.
Why would a model use Superposition?
A model has limited activation dimensions but may benefit from many potential features. If those features are sparse and seldom active together, non-orthogonal encoding can preserve more useful information than assigning one exclusive dimension per feature, at the cost of occasional interference.
What is the difference between Superposition and polysemanticity?
Superposition is a proposed representational geometry involving non-orthogonal feature directions. Polysemanticity is the observation that one neuron responds to multiple concepts or contexts. Superposition can cause polysemantic neurons, but polysemanticity alone does not establish the full geometry or its optimization cause.
How can Superposition be measured?
Researchers examine feature sparsity, dimensionality, pairwise direction overlap, interference, reconstruction error, and intervention effects. Results depend on how features are defined and recovered, so measurements should be compared across dictionary sizes, datasets, seeds, and causal tests.
Do sparse autoencoders eliminate Superposition?
Not necessarily. A sparse autoencoder attempts to express dense activations with a larger sparse dictionary, making candidate features easier to inspect. Its reconstruction can still be imperfect, latents can split or merge concepts, and the learned dictionary does not prove that the original model uses one unique decomposition.