Ten Ways Mechanistic Interpretability Research Can Mislead You
A circuit that looks clean, a feature that reconstructs well, an intervention that moves the metric — each is compatible with a story nobody has actually tried to break.
Tagged circuits · show all articles
A circuit that looks clean, a feature that reconstructs well, an intervention that moves the metric — each is compatible with a story nobody has actually tried to break.
Before any claim about what a model is doing, someone ran a forward pass, trained a probe, decomposed an activation into sparse pieces, and changed something on purpose to see what moved. This is that process, in the order it actually happens.
Long before the term existed, researchers tried to read meaning into a trained network's units. This is mechanistic interpretability's dated history: a naming paper, a landmark result, and three labs converging on one tool within months of each other.
Taking a mechanism apart tells you what each part was holding — but only once the part is off, and only for the mechanism on the bench. What interpretability has established, and what it has not.