Model Evaluation
From Origins to Frontier: A History of Frontier AI Model Comparisons
Model comparison began with judged retrieval printouts, not a leaderboard. Every mechanism since — leaderboards, broad suites, contamination audits, preference arenas, models judging models — patched a failure the last one could no longer hide.