Crosscoder
A sparse dictionary-learning variant that jointly learns shared features across multiple layers or across two models (e.g. base vs fine-tuned), enabling feature-level diffing of what training changed.
grounded in: trend theme 'Interpretability rigor: Researchers are applying formal causal methods to actually explain LLM internals'; extends the existing SAE/transcoder cluster
Connected concepts
Sparse autoencoder, Transcoder, Circuit tracing, Model migration
Explore it live in the knowledge graph →