Model Evaluation
The Hardest Unsolved Problems in AI Agent Evaluation and Reliability
A benchmark score describes a trajectory that has already ended. The hardest problems in agent evaluation are about the ones still running, the ones nobody re-ran, and the ones no aggregate number shows.