Equation 15 · Measuring AI Agent Reliability: What the Evidence Actually Supports
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the number of samples. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The two formulas share the same n , c , and k and the same combinatorial machinery, and that similarity is precisely the point: they are estimating two different probabilities from the same data, one about the best of k attempts and one about the consistency of all k of them. A system can score well on the first and badly on the second, and a report that only states pass@k while an agent is actually being deployed to run the same class of task repeatedly, unsupervised, with no way to pick the lucky run, is answering a question nobody was asking.
Sources cited in the article section
- [2] Evaluating Large Language Models Trained on Code ↗
- [3] tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains ↗
These citations give research context. Read each source to check which claims it supports.
Return to Measuring AI Agent Reliability: What the Evidence Actually Supports