raghu@dark-factory :~/kb/causal-abstraction $ cat

Causal abstraction

A mechanistic-interpretability framework that tests whether a high-level causal model of a task is actually implemented inside a network by intervening on internal activations and checking that the abstraction holds under those interventions.

grounded in: Trend theme: 'Interpretability and causal reasoning about LLMs — appetite for rigorous, mechanistic explanations of what models actually do inside.'

Connected concepts

Mechanistic interpretability, Activation patching, Circuit tracing, Chain-of-thought faithfulness

Explore it live in the knowledge graph →