raghu@dark-factory :~/kb/auto-interpretability $ cat

Automated interpretability

Using an LLM to automatically generate and score natural-language explanations for the features or neurons surfaced by interpretability tools, making feature analysis scalable beyond manual inspection.

grounded in: latest-trends theme 'LLM interpretability: Serious research effort is going into causally explaining what models do inside' — pairs the interpretability push with the system's existing llm-as-judge pa

Connected concepts

Sparse autoencoder, Dictionary learning, Monosemanticity, Mechanistic interpretability, LLM-as-judge

Explore it live in the knowledge graph →