Transcoder
An interpretability tool that replaces an MLP block with a sparse, more interpretable approximation to trace features across layers.
grounded in: latest-trends theme 'Mechanistic interpretability: Interest is rising in causal, explainable accounts of how models actually reach outputs'
Connected concepts
Sparse autoencoder, Circuit tracing, Superposition, Mechanistic interpretability
Explore it live in the knowledge graph →