Sparse autoencoder
A dictionary-learning probe that decomposes a model's dense activations into sparse, human-interpretable features to reverse-engineer what a network has learned.
grounded in: latest trends theme: 'a genuine push to understand how models actually reason' (LLM interpretability)
Connected concepts
Mechanistic interpretability, Chain-of-thought faithfulness, Hallucination control
Explore it live in the knowledge graph →