A History of Multimodal AI
Before a single model could read an image and write about it, two research communities spent a decade building separate pieces that had to be joined by hand. This is the record of each join.
Before a single model could read an image and write about it, two research communities spent a decade building separate pieces that had to be joined by hand. This is the record of each join.
A circuit that looks clean, a feature that reconstructs well, an intervention that moves the metric — each is compatible with a story nobody has actually tried to break.
A circuit that looks real on your screen has to survive a matched control, a blind check, and someone else's attempt to break it before it earns a place in an audit report. This is how the practitioners who take that seriously actually work.
Four ways of looking inside a trained network answer four different questions — what is decodable, what is the code's basis, what is causally necessary, and what can be moved by hand. None of the four subsumes the others.
Two questions decide how mechanistic interpretability matures as a technical field by 2035: whether it scales into trusted, near-complete audits, and whether the field converges on one validated toolkit. Four scenarios, each with a falsifier.
Before any claim about what a model is doing, someone ran a forward pass, trained a probe, decomposed an activation into sparse pieces, and changed something on purpose to see what moved. This is that process, in the order it actually happens.
A feature-visualization narrative that looks right and a circuit that has been checked against a known answer are different kinds of evidence. Most published claims sit closer to the first than the coverage of them usually admits.
Long before the term existed, researchers tried to read meaning into a trained network's units. This is mechanistic interpretability's dated history: a naming paper, a landmark result, and three labs converging on one tool within months of each other.
Five problems in AI alignment remain open by the researchers' own account — not by outside speculation — from judging outputs no human can fully check to deciding whose values a system should encode when people disagree.
Refusal training, content classifiers and red-team programs are not proof a deployed system is safe. They are ten separately documented ways such systems keep failing, each with its own paper trail and its own cause.