← Back to article

Equation 11 · A History of How We Learned to Evaluate AI Agents

What does this equation mean?

nn

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

the number of samples. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

nn

Symbol n

the number of samples.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

The term inside the brackets is the probability that a random draw of k items from the n samples contains no passing solution, so one minus that quantity is the probability at least one does. The point of writing the estimator this way, rather than simply sampling k times per problem, is to separate two things later agent benchmarks would have to separate again and again: how good a system is, and how much it was allowed to try. Every agent benchmark discussed below that reports a success rate is implicitly answering the question this estimator first made explicit — success at what sampling budget, counted how.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to A History of How We Learned to Evaluate AI Agents

Browse the mathematical compendium →