Published equation contexts
Why this formula appears here
That last detail forced a genuine statistical problem into agent-adjacent evaluation for the first time: naively estimating the chance that at least one of k sampled attempts succeeds, by drawing exactly k samples and checking, is a high-variance estimator, especially at small k . Chen and colleagues instead drew a larger fixed pool of n samples per problem, counted the number c that passed, and computed an unbiased estimate of the pass rate at budget k directly from that pool: . The term inside the brackets is the probability that a random draw of k items from the n samples contains no passing solution, so one minus that quantity is the probability at least one does. The…
Read the representative guide
Symbol E_problems
roblems appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →Symbol c
c occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Read this term in its guide →Denominator: binomnk
The complete quantity below the fraction bar; it must be nonzero for this division.
Read this term in its guide →How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
Research cited beside this formula
Published contexts (1)
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 9 · Model Evaluation
A History of How We Learned to Evaluate AI Agents
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
That last detail forced a genuine statistical problem into agent-adjacent evaluation for the first time: naively estimating the chance that at least one of k sampled attempts succeeds, by drawing exactly k samples and checking, is a high-variance estimator, especially at small k . Chen and colleagues instead drew a larger fixed pool of n samples per problem, counted the number c that passed, and computed an unbiased estimate of the pass rate at budget k directly from that pool: . The term inside the brackets is the probability that a random draw of k items from the n samples contains no passing solution, so one minus that quantity is the probability at least one does. The…