Feature Steering
Feature steering is an inference-time intervention that changes a model's behavior by increasing, decreasing, or otherwise controlling the activation of an identified internal feature. In the reviewed sparse-autoencoder form, a learned dictionary maps model activations to candidate features, and researchers manipulate one or more feature coefficients before continuing the forward pass. It is a subtype of activation steering, not a synonym for every steering vector or for feature discovery itself.
Origin and context
Anthropic's May 2024 Scaling Monosemanticity report extracted features from an intermediate layer of Claude 3 Sonnet and included experiments that clamped selected features to different activation strengths. The accompanying demonstration showed that amplifying a Golden Gate Bridge feature made the topic dominate unrelated responses. Independent November 2024 preprints then used sparse autoencoders to target steering vectors more precisely and to study refusal steering, including tradeoffs between stronger refusal behavior and general capabilities.
Why it matters
Prompts influence behavior through the input, while fine-tuning changes model weights. Feature steering offers a third experimental control surface inside a forward pass. Researchers can test whether a representation has a causal effect, probe entanglement among concepts, or explore temporary behavior changes without retraining the entire model. The same access can suppress desirable behavior or bypass safeguards, so the technique is best understood as a research intervention whose effect and side effects require measurement, not as a ready-made safety control.
Example
A researcher first trains or obtains a sparse autoencoder for a model layer, selects a feature associated with refusal, and increases its activation during inference. The team then compares refusal rates and unrelated benchmark performance with an unmodified baseline across several steering strengths. If refusal increases while benign-task performance falls, the experiment indicates a tradeoff rather than a clean safety switch. Simply finding that a feature correlates with refusal is not feature steering until the activation is intervened on.
How it differs
Sparse Autoencoders (SAEs)
A sparse autoencoder is a representation-learning method used to discover candidate features. Feature steering is an intervention using selected representations. An SAE can be analyzed without steering, and steering can use representations produced by other methods.
Alignment Faking
Alignment faking is a strategically conditional behavior studied across training or monitoring contexts. Feature steering modifies activations. A steering-induced behavioral change can motivate a diagnostic hypothesis but does not by itself establish strategic intent or alignment faking.
Maturity and evidence
Maturity is rated 3. The technique has a detailed primary demonstration and multiple independent implementations that compare behavioral effects and side effects. It remains below 4 because feature dictionaries are incomplete, interventions are model- and layer-specific, and robust production use has not been established.
Limits and open questions
Features may be polysemantic, incomplete, unstable across contexts, or entangled with capabilities that should remain unchanged. Steering strength can produce nonlinear and off-target behavior, and a result on one layer or model may not transfer. Access usually requires model internals, and observed control does not prove that the selected feature is the sole mechanism. Safety claims need adversarial, capability-retention, and distribution-shift evaluation.
Related terms
References
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 SonnetAnthropic / Transformer Circuits · 2024-05-21 · class A
- Improving Steering Vectors by Targeting Sparse Autoencoder FeaturesIndependent interpretability researchers / arXiv · 2024-11-04 · class A
- Steering Language Model Refusal with Sparse AutoencodersMicrosoft Research / arXiv · 2024-11-18 · class A
- Alignment faking in large language modelsAnthropic and Redwood Research / arXiv · 2024-12-18 · class A
Last updated: 2026-09-04
This term is also covered in the Skills Atlas as mechanistic interpretability skill.
This term is also covered in the Skills Atlas as ai risk management skill.