AI Agents & Systems
Measuring AI Agent Reliability: What the Evidence Actually Supports
A single successful run shows a system can succeed once. Pass@k, pass^k, and repeated-trial statistics are how anyone finds out whether it succeeds reliably — and task length changes the answer.