Causal abstraction
A mechanistic-interpretability framework that tests whether a high-level causal model of a task is actually implemented inside a network by intervening on internal activations and checking that the abstraction holds under those interventions.
grounded in: Trend theme: 'Interpretability and causal reasoning about LLMs — appetite for rigorous, mechanistic explanations of what models actually do inside.'
Connected concepts
Mechanistic interpretability, Activation patching, Circuit tracing, Chain-of-thought faithfulness
Explore it live in the knowledge graph →