raghu@dark-factory :~/kb/dictionary-learning $ cat

Dictionary learning

An unsupervised technique that decomposes a model's dense, polysemantic activations into a large overcomplete set of sparse, human-interpretable features, forming the mathematical basis for sparse autoencoders and their variants.

grounded in: latest-trends theme 'LLM interpretability: Serious research effort is going into causally explaining what models do inside'; it is the foundational method underneath the existing sparse-autoencoder/cr

Connected concepts

Sparse autoencoder, Monosemanticity, Superposition, Crosscoder, Transcoder, Mechanistic interpretability

Explore it live in the knowledge graph →