raghu@dark-factory :~/kb/logit-lens $ cat

Logit Lens

An interpretability technique that projects a transformer's intermediate-layer hidden states through the output unembedding to read the model's evolving next-token prediction across depth.

grounded in: Latest AI/tech trends, HN front page 2026-07-13 theme 'Mechanistic interpretability of LLMs: Researchers are pushing to actually explain model internals rather than treat them as black boxes'

Connected concepts

Mechanistic interpretability, Superposition, Activation patching, CoT Monitorability

Explore it live in the knowledge graph →