AI Research
Mechanistic Interpretability in 2035: Scenarios and Falsifiers
Two questions decide how mechanistic interpretability matures as a technical field by 2035: whether it scales into trusted, near-complete audits, and whether the field converges on one validated toolkit. Four scenarios, each with a falsifier.