← Back to article

Equation 11 · Measuring AI Agent Reliability: What the Evidence Actually Supports

What does this equation mean?

kk

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

the number of samples. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

kk

Symbol k

the number of samples.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Pass@k answers a specific question: given a budget of k independent attempts, what is the chance that at least one succeeds? That is the right question for a search-and-verify workflow, where a cheap checker can identify the one attempt that worked among several candidates. It is the wrong question for almost everything else an agent does, because most agentic tasks do not offer a free, cheap oracle that can pick the winning attempt out of a pile of candidates after the fact — the “attempt” is the deployment. For that setting, Yao and colleagues, building the tau-bench benchmark for tool-using agents interacting with simulated customers under domain policies, proposed the complementary…
Read the full surrounding passage
Pass@k answers a specific question: given a budget of k independent attempts, what is the chance that at least one succeeds? That is the right question for a search-and-verify workflow, where a cheap checker can identify the one attempt that worked among several candidates. It is the wrong question for almost everything else an agent does, because most agentic tasks do not offer a free, cheap oracle that can pick the winning attempt out of a pile of candidates after the fact — the “attempt” is the deployment. For that setting, Yao and colleagues, building the tau-bench benchmark for tool-using agents interacting with simulated customers under domain policies, proposed the complementary statistic: the probability that every one of k independent trials on the same task succeeds, rather than at least one of them [ 3 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Measuring AI Agent Reliability: What the Evidence Actually Supports

Browse the mathematical compendium →