raghu@dark-factory :~/kb/sparse-autoencoder $ cat

Sparse autoencoder

A dictionary-learning probe that decomposes a model's dense activations into sparse, human-interpretable features to reverse-engineer what a network has learned.

grounded in: latest trends theme: 'a genuine push to understand how models actually reason' (LLM interpretability)

Connected concepts

Mechanistic interpretability, Chain-of-thought faithfulness, Hallucination control

Explore it live in the knowledge graph →