What is Patchscopes?

Patchscopes are configurable interpretability experiments that extract a hidden representation from a source model run, optionally transform it, insert it at a selected position and layer in a target run, and decode the resulting behavior.

Quick Facts

SpecificationOfficial Specification

How It Works

Specify the source-to-target configuration

Record the source model and revision, source prompt, layer, token position, activation type, transformation, target model, target prompt, target layer, target position, and decoding rule. The Patchscopes paper presents this modular configuration and shows how vocabulary projection, feature extraction, entity description, and cross-model inspection can be expressed as related patching setups.

Patch and decode under controlled prompts

Run the source model to capture the selected hidden state, transform dimensions or bases only when required, run the target context to its insertion point, replace the target state, and continue generation. Compare multiple target prompts, source and target layers, positions, temperatures, and decoding constraints. The official project page documents token-identity, attribute, entity, and cross-model examples.

Test whether the activation adds privileged evidence

Use no-patch, zero-vector, shuffled-activation, input-only, and mismatched-knowledge controls. Recent evaluation of activation verbalization found that a verbalizer can answer from the visible prompt or its own parametric knowledge rather than the target activation. Therefore measure incremental information over controls, robustness to prompt changes, source-target compatibility, and downstream causal relevance before treating generated text as an explanation.

Key Characteristics

  • Moves a selected hidden representation from a source run into a target run
  • Allows optional transformations between source and target representation spaces
  • Uses target prompts and model generation as configurable decoding interfaces
  • Can inspect token identity, attributes, contextualization, or cross-model mappings
  • Requires baselines that separate activation evidence from prompt and model priors
  • Produces hypothesis-generating descriptions rather than guaranteed faithful explanations

Common Use Cases

  1. Decoding token identity from hidden states at different model layers
  2. Verbalizing candidate entity information represented before the final layer
  3. Comparing how a concept is represented across compatible model checkpoints
  4. Testing whether a stronger target model can decode a smaller source model's state
  5. Auditing activation verbalization against no-patch and mismatched-knowledge controls

Example

loading...
Loading code...

Frequently Asked Questions

What information defines a Patchscope?

A reproducible Patchscope identifies the source model, prompt, layer, position, and activation; any source-to-target transformation; the target model, prompt, layer, and position; and the decoding and evaluation rules. Omitting one of these can change the intervention being interpreted.

How are Patchscopes different from Activation Patching?

Activation Patching usually asks whether a replacement restores or changes a task metric in a matched run. Patchscopes use a separately designed target context as a decoding interface and can patch across prompts or models. Localization is one possible configuration, not the framework's only purpose.

Do Patchscopes require a trained probe?

Not necessarily. Same-model configurations can use an identity transformation and a carefully designed target prompt. Cross-model or dimension-mismatched configurations may require a learned mapping. Either way, target-prompt sensitivity and target-model priors must be measured on held-out controls.

Can a Patchscope verbalization be unfaithful?

Yes. The target model may infer an answer from visible prompt text or supply knowledge from its own parameters. Compare against no-patch, zero, shuffled, and input-only baselines, and create cases where source and target knowledge disagree. Fluency and task accuracy alone do not establish privileged access.

Can Patchscopes prove what a source model is thinking?

No. They reveal what a configured target computation can decode after receiving a source activation. This supports a representation hypothesis under the tested setup, not a complete account of internal reasoning. Causal interventions and behavior-level tests are still required for stronger claims.

Related Terms