Circuit tracing
Reconstructing the internal computation a model uses for a specific behavior by building attribution graphs over its learned features, to explain — not merely observe — how it reasons.
grounded in: trend theme 'Can we actually understand LLM reasoning: Interpretability of how models reason is a recurring serious-research thread'; deepens the existing interp cluster (mechanistic-interpretability,
Connected concepts
Mechanistic interpretability, Sparse autoencoder, Chain-of-thought faithfulness
Explore it live in the knowledge graph →