What is Activation Steering?
Activation Steering is an inference-time intervention that modifies a model's internal activations along a selected direction or subspace to increase, decrease, or otherwise control a measured behavior without updating model weights.
Quick Facts
| Specification | Official Specification |
|---|
How It Works
Estimate a behavior-relevant direction
A common contrastive method collects activations from matched positive and negative examples, then subtracts their means. Other methods use linear probes, principal directions, sparse-autoencoder features, or optimized vectors. Train and evaluate on separated prompt sets, control lexical and formatting cues, and test whether the direction predicts the intended property rather than dataset artifacts.
Apply the intervention at inference time
Choose the model component, layer, token positions, schedule, normalization, and coefficient. Activation Addition demonstrates adding contrast-derived vectors to transformer activations without changing weights, while Representation Engineering studies population-level representations for monitoring and control. Layer and strength sweeps are empirical requirements, not cosmetic tuning.
Evaluate target effects and collateral changes
Measure the intended behavior on held-out and adversarial prompts, then assess task quality, refusal errors, factuality, calibration, style leakage, multilingual transfer, latency, and interactions with sampling. Compare against prompting, decoding controls, and fine-tuning under the same evaluation. A steering vector can expose or influence a representation without proving that it is necessary, sufficient, disentangled, or safe.
Key Characteristics
- Changes internal activations during inference without updating model weights
- Uses a selected direction, feature, or subspace tied to a target behavior
- Requires explicit layer, token-position, schedule, and strength choices
- Can be estimated from contrastive data, probes, or learned feature dictionaries
- Produces prompt- and context-dependent effects with possible collateral changes
- Requires behavioral, quality, robustness, and safety evaluation before deployment
Common Use Cases
- Testing whether an internal representation can influence a target behavior
- Adjusting style or behavioral tendencies in controlled model experiments
- Comparing inference-time control with prompting or weight fine-tuning
- Probing trade-offs between target gains and unrelated capability changes
- Building research prototypes for representation monitoring and control
Example
Loading code...Frequently Asked Questions
How is Activation Steering different from fine-tuning?
Fine-tuning changes model parameters through optimization and produces a new checkpoint or adapter. Activation Steering leaves weights unchanged and modifies selected internal states during inference. Steering can be fast to test, but it adds runtime intervention logic and may be less stable across prompts.
How is Activation Steering different from Activation Patching?
Patching transfers activations from a controlled comparison run, usually to test causal localization. Steering applies a chosen direction or subspace to influence behavior, often across many new inputs. The first is primarily an explanatory intervention; the second is primarily a control intervention, though experiments can overlap.
How is an activation steering vector created?
A common approach subtracts mean activations from matched negative examples from those of positive examples at a selected layer. Alternatives include probe weights, principal directions, optimized vectors, and sparse-autoencoder features. Held-out tests must verify that the vector captures the target rather than superficial cues.
Where and how strongly should a steering vector be applied?
There is no universally correct layer or coefficient. Sweep layers, token positions, generation steps, and strengths on validation data. Record the intervention convention and activation scale, then choose an operating point using both target improvement and collateral quality or safety metrics.
Does Activation Steering guarantee aligned or safe behavior?
No. Steering can fail on distribution shifts, be overridden by context, or change unrelated capabilities. A direction associated with a safety-relevant behavior is not proof of a complete causal mechanism. Production use still needs adversarial evaluation, monitoring, access controls, and fallback policies.