← All parts of this equation

Equation 7 · Part 2 · Measuring AI Agent Reliability: What the Evidence Actually Supports

Symbol E_tasks

pass@k  :=  Etasks ⁣[ 1  −  (n−ck)(nk) ]\text{pass@}k \;:=\; \mathbb{E}_{\text{tasks}}\!\left[\, 1 \;-\; \frac{\binom{n-c}{k}}{\binom{n}{k}} \,\right]
Etasks\mathbb{E}_{\text{tasks}}

What this part means

EtE_tasks appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Its job in the formula

EtE_tasks appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

The passage around this formula

Start with the metric that made repeated sampling a standard practice. When OpenAI’s Codex team evaluated a code-generating model against the HumanEval benchmark, they needed a way to score the strategy of drawing several candidate solutions from the model and keeping the best one. The obvious approach — generate exactly k samples per problem and check whether any of them pass — has an undesirable property: it is a valid estimate but a high-variance one, since it throws away information every time you happen to generate more or fewer than k samples. Their fix was to over-sample: draw n total samples per task, observe how many of them, c , actually pass, and then compute the exact probability…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.