What is Causal Scrubbing?

Causal Scrubbing is a mechanistic interpretability method that converts an explicit computational hypothesis into resampling interventions and tests whether the model preserves relevant behavior when internal activations are replaced according to equivalence relations licensed by that hypothesis.

Quick Facts

SpecificationOfficial Specification

How It Works

Write an explicit correspondence before testing

State the abstract variables, dependencies, model components, token positions, input distribution, and behavior metric. Define which inputs are equivalent for each abstract node and which concrete activation implements it. The Causal Scrubbing method requires interventions to follow this interpretation rather than choosing convenient donors after seeing results.

Run hypothesis-licensed resampling ablations

For each tested node, sample donor inputs that agree on the abstract information the hypothesis says must be preserved, then replace the mapped activation and continue the model computation. Repeat over donors, examples, and random seeds. Compare the scrubbed behavior with the original model, unrestricted resampling, random or mean ablation, and alternative hypotheses under the same metric.

Interpret preservation and degradation asymmetrically

If behavior degrades substantially, the tested correspondence is insufficient under that intervention or the implementation is wrong. If behavior is preserved, the result is consistent with the hypothesis but may also fit coarser, redundant, or alternative mechanisms. Conclusions depend on donor support, equivalence-class quality, metric sensitivity, and whether resampled activations stay on a realistic distribution.

Key Characteristics

  • Starts from an explicit abstract computational hypothesis and model correspondence
  • Defines input equivalence classes that license internal resampling interventions
  • Replaces concrete activations with values from hypothesis-compatible donor inputs
  • Measures preservation or degradation on a predeclared behavior metric
  • Requires donor, ablation, alternative-hypothesis, and distribution controls
  • Can falsify or support a tested explanation but cannot establish uniqueness

Common Use Cases

  1. Testing whether a proposed circuit computes only the abstract variables it claims
  2. Checking whether model components are interchangeable within hypothesis-defined classes
  3. Comparing competing mechanistic explanations under matched interventions
  4. Finding missing dependencies when a scrubbed model loses target behavior
  5. Turning informal circuit diagrams into reproducible intervention contracts

Example

loading...
Loading code...

Frequently Asked Questions

What does passing a Causal Scrubbing test prove?

It shows that the model preserved the selected behavior under the resampling interventions generated by one hypothesis, interpretation, dataset, and metric. This is evidence consistent with that explanation, but it does not prove the explanation is unique, minimal, complete, or valid outside the tested distribution.

How is Causal Scrubbing different from Activation Patching?

Activation Patching often tests whether replacing one activation from a clean run restores a corrupted metric. Causal Scrubbing starts with an explicit abstract hypothesis and systematically chooses donors through its equivalence relations. Patching is an intervention primitive; scrubbing is a hypothesis-testing protocol built from resampling interventions.

How are donor inputs selected in Causal Scrubbing?

Donors are sampled from inputs that agree on the abstract information a tested node is supposed to represent, including parent-dependent constraints where the hypothesis requires them. The rule must be fixed before results are inspected. Sparse or biased equivalence classes weaken the test and should be reported.

Why can a correct hypothesis fail a Causal Scrubbing test?

The concrete mapping may be wrong, the resampled activation may be off distribution, the metric may include unrelated behavior, or the hypothesis may omit a dependency needed by the implementation. Failures therefore reject the tested correspondence as a whole and guide refinement; they do not identify the broken assumption automatically.

Is Causal Scrubbing the same as causal inference with observational data?

No. It performs controlled internal resampling interventions on a known model computation to test a mechanistic abstraction. Classical causal inference often estimates effects in real-world systems with partial observability, treatment assignment, and identification assumptions. The fields share causal language but answer different questions.

Related Terms