raghu@dark-factory :~/kb/circuit-tracing $ cat

Circuit tracing

Reconstructing the internal computation a model uses for a specific behavior by building attribution graphs over its learned features, to explain — not merely observe — how it reasons.

grounded in: trend theme 'Can we actually understand LLM reasoning: Interpretability of how models reason is a recurring serious-research thread'; deepens the existing interp cluster (mechanistic-interpretability,

Connected concepts

Mechanistic interpretability, Sparse autoencoder, Chain-of-thought faithfulness

Explore it live in the knowledge graph →