Equation 3 · A Practitioner's Map of Agent Evaluation Frameworks
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol hat p
hat p occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Symbol z_α/2
z_α/2 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol n
n occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction.
What the article says around this equation
The second argument is about how much any one number can actually tell you, independent of harness effects, purely from how many tasks produced it. Treat a benchmark’s headline pass rate as an estimate of a task-set success probability, drawn from n roughly independent trials. The familiar normal approximation to a binomial confidence interval, . gives a rough sense of how much noise sits under a given n , even granting the generous and almost certainly false assumption that every task in the set is an equally difficult coin flip. At the widest plausible spread, = 0.5 , a 95% interval ( 1.96 ) works out to roughly 13.9 points on a…
Read the full surrounding passage
The second argument is about how much any one number can actually tell you, independent of harness effects, purely from how many tasks produced it. Treat a benchmark’s headline pass rate as an estimate of a task-set success probability, drawn from n roughly independent trials. The familiar normal approximation to a binomial confidence interval, . gives a rough sense of how much noise sits under a given n , even granting the generous and almost certainly false assumption that every task in the set is an equally difficult coin flip. At the widest plausible spread, = 0.5 , a 95% interval ( 1.96 ) works out to roughly 13.9 points on a fifty-task set — close to the size of tau-bench’s original airline domain or GAIA’s seventy-five-question Level 3 tier — versus roughly 4.4 points on SWE-bench Verified’s five hundred tasks, or 5.1 points on OSWorld’s 369. A three- or four-point difference between two systems on a fifty- or seventy-five-task slice is, on this generous accounting, close to noise; the same gap on a five-hundred-task set is not. Real task-to-task difficulty variation only widens these intervals further, so treat the numbers above as a floor on the uncertainty, not the whole of it. The lesson is not to distrust small benchmarks — tau-bench’s airline domain and GAIA’s Level 3 tier are deliberately small because hard, well-verified tasks in those domains are expensive to build — but to stop reading a two-point movement on either of them as a finding.
Sources cited in the article section
These citations give research context. Read each source to check which claims it supports.
Return to A Practitioner's Map of Agent Evaluation Frameworks