Activation steering
Controlling model behavior at inference by adding interpretable feature directions into the residual stream instead of retraining or prompting.
grounded in: latest trends theme: interpretability push to understand and intervene on how models reason
Connected concepts
Mechanistic interpretability, Sparse autoencoder, Guardrails / bounded autonomy
Explore it live in the knowledge graph →