Dictionary learning
An unsupervised technique that decomposes a model's dense, polysemantic activations into a large overcomplete set of sparse, human-interpretable features, forming the mathematical basis for sparse autoencoders and their variants.
grounded in: latest-trends theme 'LLM interpretability: Serious research effort is going into causally explaining what models do inside'; it is the foundational method underneath the existing sparse-autoencoder/cr
Connected concepts
Sparse autoencoder, Monosemanticity, Superposition, Crosscoder, Transcoder, Mechanistic interpretability
Explore it live in the knowledge graph →