← Back to article

Equation 12 · Measuring AI Agent Reliability: What the Evidence Actually Supports

What does this equation mean?

passk  :=  Etasks ⁣[ (ck)(nk) ]\text{pass}^{k} \;:=\; \mathbb{E}_{\text{tasks}}\!\left[\, \frac{\binom{c}{k}}{\binom{n}{k}} \,\right]

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Start withbinomck
Divide bybinomnk
This relates topass^k :
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

kk

Symbol k

the number of samples.

Understand this part →

Etasks\mathbb{E}_{\text{tasks}}

Symbol E_tasks

EtE_tasks appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

cc

Symbol c

c occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

nn

Symbol n

the two formulas share the same.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →
(ck)\binom{c}{k}

Numerator: binomck

The complete quantity above the fraction bar.

Understand this part →

(nk)\binom{n}{k}

Denominator: binomnk

The complete quantity below the fraction bar; it must be nonzero for this division.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Pass@k answers a specific question: given a budget of k independent attempts, what is the chance that at least one succeeds? That is the right question for a search-and-verify workflow, where a cheap checker can identify the one attempt that worked among several candidates. It is the wrong question for almost everything else an agent does, because most agentic tasks do not offer a free, cheap oracle that can pick the winning attempt out of a pile of candidates after the fact — the “attempt” is the deployment. For that setting, Yao and colleagues, building the tau-bench benchmark for tool-using agents interacting with simulated customers under domain policies, proposed the complementary…
Read the full surrounding passage
Pass@k answers a specific question: given a budget of k independent attempts, what is the chance that at least one succeeds? That is the right question for a search-and-verify workflow, where a cheap checker can identify the one attempt that worked among several candidates. It is the wrong question for almost everything else an agent does, because most agentic tasks do not offer a free, cheap oracle that can pick the winning attempt out of a pile of candidates after the fact — the “attempt” is the deployment. For that setting, Yao and colleagues, building the tau-bench benchmark for tool-using agents interacting with simulated customers under domain policies, proposed the complementary statistic: the probability that every one of k independent trials on the same task succeeds, rather than at least one of them [ 3 ] . The two formulas share the same n , c , and k and the same combinatorial machinery, and that similarity is precisely the point: they are estimating two different probabilities from the same data, one about the best of k attempts and one about the consistency of all k of them. A system can score well on the first and badly on the second, and a report that only states pass@k while an agent is actually being deployed to run the same class of task repeatedly, unsupervised, with no way to pick the lucky run, is answering a question nobody was asking.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Measuring AI Agent Reliability: What the Evidence Actually Supports

See this formula across 1 published context →

Browse the mathematical compendium →