raghu@dark-factory :~/kb/activation-steering $ cat

Activation steering

Controlling model behavior at inference by adding interpretable feature directions into the residual stream instead of retraining or prompting.

grounded in: latest trends theme: interpretability push to understand and intervene on how models reason

Connected concepts

Mechanistic interpretability, Sparse autoencoder, Guardrails / bounded autonomy

Explore it live in the knowledge graph →