Attribution patching
A gradient-based linear approximation of activation patching that estimates every component's causal effect in a single backward pass, making circuit attribution scale to large models.
grounded in: Trend theme 'shifting toward causal explanations of model behavior'; a durable interpretability method distinct from the activation-/path-patching concepts already present.
Connected concepts
Activation patching, Path patching, Circuit tracing, Mechanistic interpretability
Explore it live in the knowledge graph →