raghu@dark-factory :~/kb/transcoder $ cat

Transcoder

An interpretability tool that replaces an MLP block with a sparse, more interpretable approximation to trace features across layers.

grounded in: latest-trends theme 'Mechanistic interpretability: Interest is rising in causal, explainable accounts of how models actually reach outputs'

Connected concepts

Sparse autoencoder, Circuit tracing, Superposition, Mechanistic interpretability

Explore it live in the knowledge graph →