raghu@dark-factory :~/kb/activation-patching $ cat

Activation patching

A causal interpretability method that copies activations from one forward pass into another (ablate/patch) to measure which internal components are responsible for a specific model behavior.

grounded in: HN trend theme 'Interpretability & research rigor: Serious effort is going into causal, mechanistic understanding of what models actually do internally.'

Connected concepts

Mechanistic interpretability, Circuit tracing, Sparse autoencoder, Superposition

Explore it live in the knowledge graph →