A History of How We Learned to Evaluate AI Agents
Grading a model once meant scoring a string against a reference answer. Grading an agent required borrowing a games lab's tools, discarding them, and rebuilding evaluation from scratch, in a documented, dated sequence of its own.