raghu@dark-factory :~/kb/superposition $ cat

Superposition

The phenomenon whereby a neural network represents more features than it has neurons by encoding them as overlapping directions in activation space, which sparse autoencoders aim to disentangle.

grounded in: Trend theme 'Can we actually understand LLM reasoning' (interpretability research), deepening the existing sparse-autoencoder/mechanistic-interpretability cluster.

Connected concepts

Sparse autoencoder, Mechanistic interpretability, Circuit tracing

Explore it live in the knowledge graph →