raghu@dark-factory :~/kb/linear-probing $ cat

Linear probing

An interpretability technique that trains a simple linear classifier on a model's internal activations to test whether a concept or property is linearly represented at a given layer.

grounded in: Trend theme 'Mechanistic interpretability: opening the black box of LLMs with rigorous causal methods is a rising research focus'; a foundational probing method absent from the graph's SAE/steering/pa

Connected concepts

Mechanistic interpretability, Logit Lens, Sparse autoencoder, Activation steering

Explore it live in the knowledge graph →