← Mathematical compendium

Published equation contexts

passk  :=  Etasks ⁣[ (ck)(nk) ]\text{pass}^{k} \;:=\; \mathbb{E}_{\text{tasks}}\!\left[\, \frac{\binom{c}{k}}{\binom{n}{k}} \,\right]

Why this formula appears here

Pass@k answers a specific question: given a budget of k independent attempts, what is the chance that at least one succeeds? That is the right question for a search-and-verify workflow, where a cheap checker can identify the one attempt that worked among several candidates. It is the wrong question for almost everything else an agent does, because most agentic tasks do not offer a free, cheap oracle that can pick the winning attempt out of a pile of candidates after the fact — the “attempt” is the deployment. For that setting, Yao and colleagues, building the tau-bench benchmark for tool-using agents interacting with simulated customers under domain policies, proposed the complementary…

Read the full article-specific guide →

Read the representative guide

Etasks\mathbb{E}_{\text{tasks}}

Symbol E_tasks

EtE_tasks appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
(nk)\binom{n}{k}

Denominator: binomnk

The complete quantity below the fraction bar; it must be nonzero for this division.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

passk  :=  Etasks ⁣[ (ck)(nk) ]\text{pass}^{k} \;:=\; \mathbb{E}_{\text{tasks}}\!\left[\, \frac{\binom{c}{k}}{\binom{n}{k}} \,\right]

Equation 12 · AI Agents & Systems

Measuring AI Agent Reliability: What the Evidence Actually Supports

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Pass@k answers a specific question: given a budget of k independent attempts, what is the chance that at least one succeeds? That is the right question for a search-and-verify workflow, where a cheap checker can identify the one attempt that worked among several candidates. It is the wrong question for almost everything else an agent does, because most agentic tasks do not offer a free, cheap oracle that can pick the winning attempt out of a pile of candidates after the fact — the “attempt” is the deployment. For that setting, Yao and colleagues, building the tau-bench benchmark for tool-using agents interacting with simulated customers under domain policies, proposed the complementary…

Meanings in this article

  • kk: the number of samples.
  • nn: the two formulas share the same.
Equation guide → · Article →