AI Research
Probing, Sparse Autoencoders, Patching, and Steering: The Main Interpretability Methods, Compared
Four ways of looking inside a trained network answer four different questions — what is decodable, what is the code's basis, what is causally necessary, and what can be moved by hand. None of the four subsumes the others.