Published equation contexts
Why this formula appears here
Start with the metric that made repeated sampling a standard practice. When OpenAI’s Codex team evaluated a code-generating model against the HumanEval benchmark, they needed a way to score the strategy of drawing several candidate solutions from the model and keeping the best one. The obvious approach — generate exactly k samples per problem and check whether any of them pass — has an undesirable property: it is a valid estimate but a high-variance one, since it throws away information every time you happen to generate more or fewer than k samples. Their fix was to over-sample: draw n total samples per task, observe how many of them, c , actually pass, and then compute the exact probability…
Read the representative guide
Symbol E_tasks
asks appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →Symbol n
n appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →Symbol c
c occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Read this term in its guide →Denominator: binomnk
The complete quantity below the fraction bar; it must be nonzero for this division.
Read this term in its guide →How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
Research cited beside this formula
Published contexts (1)
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 7 · AI Agents & Systems
Measuring AI Agent Reliability: What the Evidence Actually Supports
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
Start with the metric that made repeated sampling a standard practice. When OpenAI’s Codex team evaluated a code-generating model against the HumanEval benchmark, they needed a way to score the strategy of drawing several candidate solutions from the model and keeping the best one. The obvious approach — generate exactly k samples per problem and check whether any of them pass — has an undesirable property: it is a valid estimate but a high-variance one, since it throws away information every time you happen to generate more or fewer than k samples. Their fix was to over-sample: draw n total samples per task, observe how many of them, c , actually pass, and then compute the exact probability…