raghu@dark-factory :~/kb/monosemanticity $ cat

Monosemanticity

The interpretability goal of decomposing a model's activations into features that each carry a single, human-interpretable meaning, the property sparse autoencoders and transcoders are built to recover.

grounded in: trend theme 'Interpretability rigor: Researchers are applying formal causal methods to actually explain LLM internals' — the direct complement to the existing `superposition` concept, which the graph

Connected concepts

Superposition, Sparse autoencoder, Transcoder, Mechanistic interpretability

Explore it live in the knowledge graph →