Ten Ways Mechanistic Interpretability Research Can Mislead You
A circuit that looks clean, a feature that reconstructs well, an intervention that moves the metric — each is compatible with a story nobody has actually tried to break.
Filtered to AI Research · show all articles
A circuit that looks clean, a feature that reconstructs well, an intervention that moves the metric — each is compatible with a story nobody has actually tried to break.
A circuit that looks real on your screen has to survive a matched control, a blind check, and someone else's attempt to break it before it earns a place in an audit report. This is how the practitioners who take that seriously actually work.
Four ways of looking inside a trained network answer four different questions — what is decodable, what is the code's basis, what is causally necessary, and what can be moved by hand. None of the four subsumes the others.
Two questions decide how mechanistic interpretability matures as a technical field by 2035: whether it scales into trusted, near-complete audits, and whether the field converges on one validated toolkit. Four scenarios, each with a falsifier.
Before any claim about what a model is doing, someone ran a forward pass, trained a probe, decomposed an activation into sparse pieces, and changed something on purpose to see what moved. This is that process, in the order it actually happens.
A feature-visualization narrative that looks right and a circuit that has been checked against a known answer are different kinds of evidence. Most published claims sit closer to the first than the coverage of them usually admits.
Long before the term existed, researchers tried to read meaning into a trained network's units. This is mechanistic interpretability's dated history: a naming paper, a landmark result, and three labs converging on one tool within months of each other.