raghu@dark-factory :~/kb/attribution-patching $ cat

Attribution patching

A gradient-based linear approximation of activation patching that estimates every component's causal effect in a single backward pass, making circuit attribution scale to large models.

grounded in: Trend theme 'shifting toward causal explanations of model behavior'; a durable interpretability method distinct from the activation-/path-patching concepts already present.

Connected concepts

Activation patching, Path patching, Circuit tracing, Mechanistic interpretability

Explore it live in the knowledge graph →