Linear probing
An interpretability technique that trains a simple linear classifier on a model's internal activations to test whether a concept or property is linearly represented at a given layer.
grounded in: Trend theme 'Mechanistic interpretability: opening the black box of LLMs with rigorous causal methods is a rising research focus'; a foundational probing method absent from the graph's SAE/steering/pa
Connected concepts
Mechanistic interpretability, Logit Lens, Sparse autoencoder, Activation steering
Explore it live in the knowledge graph →