Equation 9 · A History of How We Learned to Evaluate AI Agents
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol E_problems
roblems appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol c
c occurs above the fraction bar. The numerator is divided by the entire denominator below it.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
Denominator: binomnk
The complete quantity below the fraction bar; it must be nonzero for this division.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
That last detail forced a genuine statistical problem into agent-adjacent evaluation for the first time: naively estimating the chance that at least one of k sampled attempts succeeds, by drawing exactly k samples and checking, is a high-variance estimator, especially at small k . Chen and colleagues instead drew a larger fixed pool of n samples per problem, counted the number c that passed, and computed an unbiased estimate of the pass rate at budget k directly from that pool: . The term inside the brackets is the probability that a random draw of k items from the n samples contains no passing solution, so one minus that quantity is the probability at least one does. The…
Read the full surrounding passage
That last detail forced a genuine statistical problem into agent-adjacent evaluation for the first time: naively estimating the chance that at least one of k sampled attempts succeeds, by drawing exactly k samples and checking, is a high-variance estimator, especially at small k . Chen and colleagues instead drew a larger fixed pool of n samples per problem, counted the number c that passed, and computed an unbiased estimate of the pass rate at budget k directly from that pool: . The term inside the brackets is the probability that a random draw of k items from the n samples contains no passing solution, so one minus that quantity is the probability at least one does. The point of writing the estimator this way, rather than simply sampling k times per problem, is to separate two things later agent benchmarks would have to separate again and again: how good a system is, and how much it was allowed to try. Every agent benchmark discussed below that reports a success rate is implicitly answering the question this estimator first made explicit — success at what sampling budget, counted how.
Sources cited in the article section
These citations give research context. Read each source to check which claims it supports.
Return to A History of How We Learned to Evaluate AI Agents