Equation 12 · A History of How We Learned to Evaluate AI Agents
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the especially at small. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The term inside the brackets is the probability that a random draw of k items from the n samples contains no passing solution, so one minus that quantity is the probability at least one does. The point of writing the estimator this way, rather than simply sampling k times per problem, is to separate two things later agent benchmarks would have to separate again and again: how good a system is, and how much it was allowed to try. Every agent benchmark discussed below that reports a success rate is implicitly answering the question this estimator first made explicit — success at what sampling budget, counted how.
Sources cited in the article section
These citations give research context. Read each source to check which claims it supports.