Superposition
The phenomenon whereby a neural network represents more features than it has neurons by encoding them as overlapping directions in activation space, which sparse autoencoders aim to disentangle.
grounded in: Trend theme 'Can we actually understand LLM reasoning' (interpretability research), deepening the existing sparse-autoencoder/mechanistic-interpretability cluster.
Connected concepts
Sparse autoencoder, Mechanistic interpretability, Circuit tracing
Explore it live in the knowledge graph →