Activation patching
A causal interpretability method that copies activations from one forward pass into another (ablate/patch) to measure which internal components are responsible for a specific model behavior.
grounded in: HN trend theme 'Interpretability & research rigor: Serious effort is going into causal, mechanistic understanding of what models actually do internally.'
Connected concepts
Mechanistic interpretability, Circuit tracing, Sparse autoencoder, Superposition
Explore it live in the knowledge graph →