What is Mechanistic Interpretability?
Mechanistic Interpretability is the study of how a trained model implements a computation by identifying internal representations, components, and interactions, then testing whether the proposed mechanism causally explains specified behavior.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Move from model parts to candidate circuits
Analysis may begin with residual-stream directions, attention heads, MLP outputs, individual neurons, or features recovered by a sparse dictionary. Attribution and visualization identify components associated with a behavior, while circuit hypotheses state how those components interact across token positions and layers. The Transformer Circuits framework develops one influential vocabulary for studying these computations.
Use interventions to test causal claims
Ablation removes or replaces a component, activation patching transfers an internal state between controlled runs, and counterfactual inputs test whether the proposed computation tracks the relevant variable. Good experiments compare multiple baselines, preserve unrelated model state, predefine an output metric, and include negative controls. Localization shows where an intervention matters; it does not by itself recover the minimal or complete algorithm.
State the scope of every explanation
Mechanistic findings depend on the checkpoint, architecture, prompt distribution, token position, intervention, and metric. Features may be distributed or represented in superposition, and apparently coherent units can split, merge, or change across contexts. Report the tested behavior, data, component granularity, causal effect, uncertainty, and failed cases. Interpretability evidence can support debugging and risk analysis, but it is not a proof of model safety or intent.
Key Characteristics
- Explains specified behavior through internal representations and computations
- Builds circuit hypotheses from activations, weights, features, and information flow
- Distinguishes correlational localization from causal intervention evidence
- Scopes conclusions to a checkpoint, prompt distribution, position, and metric
- Tests robustness with baselines, controls, counterfactuals, and replication
- Does not assume that neurons, heads, or learned features are automatically human concepts
Common Use Cases
- Tracing which internal computations support a factual or algorithmic behavior
- Testing whether a suspected feature or circuit causes an unwanted output
- Comparing how related checkpoints implement the same task
- Diagnosing brittle shortcuts, context dependence, or unexpected feature interactions
- Generating auditable hypotheses for model editing, monitoring, or safety evaluation
Example
Loading code...Frequently Asked Questions
How is Mechanistic Interpretability different from explainable AI?
Explainable AI includes output-level explanations, feature attribution, examples, surrogate models, and human-facing rationales. Mechanistic Interpretability focuses more narrowly on internal representations and computations, seeking a causally tested account of how model components implement a particular behavior.
Does finding a correlated neuron explain a model behavior?
No. Correlation can identify a useful candidate, but the neuron may track a consequence, share information with other components, or matter only in one context. A stronger explanation combines localization with controlled interventions, alternative baselines, negative controls, and replication over representative inputs.
What is a circuit in Mechanistic Interpretability?
A circuit is a proposed collection of model components and interactions that implements a defined computation. Depending on the analysis, components may be attention heads, MLPs, residual-stream directions, neurons, or learned features. The circuit should predict intervention outcomes, not merely summarize a visualization.
Can Mechanistic Interpretability prove that a model is safe?
No. An analysis covers selected behaviors, prompts, components, and metrics, while untested mechanisms or distribution shifts may remain. Mechanistic evidence can reveal risks, improve monitoring, and test safety hypotheses, but it must be combined with behavioral evaluation, threat modeling, and deployment controls.
How should a mechanistic explanation be validated?
Define the behavior and metric first, locate candidate components, intervene on them, compare multiple plausible baselines, include negative controls, and test held-out prompt families or checkpoints. Report effect sizes, residual unexplained behavior, sensitivity to design choices, and cases where the mechanism fails.