Monosemanticity
The interpretability goal of decomposing a model's activations into features that each carry a single, human-interpretable meaning, the property sparse autoencoders and transcoders are built to recover.
grounded in: trend theme 'Interpretability rigor: Researchers are applying formal causal methods to actually explain LLM internals' — the direct complement to the existing `superposition` concept, which the graph
Connected concepts
Superposition, Sparse autoencoder, Transcoder, Mechanistic interpretability
Explore it live in the knowledge graph →