Automated interpretability
Using an LLM to automatically generate and score natural-language explanations for the features or neurons surfaced by interpretability tools, making feature analysis scalable beyond manual inspection.
grounded in: latest-trends theme 'LLM interpretability: Serious research effort is going into causally explaining what models do inside' — pairs the interpretability push with the system's existing llm-as-judge pa
Connected concepts
Sparse autoencoder, Dictionary learning, Monosemanticity, Mechanistic interpretability, LLM-as-judge
Explore it live in the knowledge graph →