← Back to article

Equation 3 · A Practitioner's Map of Agent Evaluation Frameworks

What does this equation mean?

p^  ±  zα/2p^(1−p^)n,\hat p \; \pm \; z_{\alpha/2}\sqrt{\frac{\hat p (1-\hat p)}{n}},

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

p^\hat p

Symbol hat p

hat p occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

zα/2z_{\alpha/2}

Symbol z_α/2

z_α/2 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

nn

Symbol n

n occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Understand this part →

fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
√

√

Take a square root.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

p^(1−p^)\hat p (1-\hat p)

Numerator: hat p (1-hat p)

The complete quantity above the fraction bar.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction.

What the article says around this equation

The second argument is about how much any one number can actually tell you, independent of harness effects, purely from how many tasks produced it. Treat a benchmark’s headline pass rate as an estimate p^\hat p of a task-set success probability, drawn from n roughly independent trials. The familiar normal approximation to a binomial confidence interval, p^  ±  zα/2p^(1−p^)n\hat p \; \pm \; z_{\alpha/2}\sqrt{\frac{\hat p (1-\hat p)}{n}}. gives a rough sense of how much noise sits under a given n , even granting the generous and almost certainly false assumption that every task in the set is an equally difficult coin flip. At the widest plausible spread, p^\hat p = 0.5 , a 95% interval ( zα/2z_{\alpha/2} ≈\approx 1.96 ) works out to roughly ±\pm 13.9 points on a…
Read the full surrounding passage
The second argument is about how much any one number can actually tell you, independent of harness effects, purely from how many tasks produced it. Treat a benchmark’s headline pass rate as an estimate p^\hat p of a task-set success probability, drawn from n roughly independent trials. The familiar normal approximation to a binomial confidence interval, p^  ±  zα/2p^(1−p^)n\hat p \; \pm \; z_{\alpha/2}\sqrt{\frac{\hat p (1-\hat p)}{n}}. gives a rough sense of how much noise sits under a given n , even granting the generous and almost certainly false assumption that every task in the set is an equally difficult coin flip. At the widest plausible spread, p^\hat p = 0.5 , a 95% interval ( zα/2z_{\alpha/2} ≈\approx 1.96 ) works out to roughly ±\pm 13.9 points on a fifty-task set — close to the size of tau-bench’s original airline domain or GAIA’s seventy-five-question Level 3 tier — versus roughly ±\pm 4.4 points on SWE-bench Verified’s five hundred tasks, or ±\pm 5.1 points on OSWorld’s 369. A three- or four-point difference between two systems on a fifty- or seventy-five-task slice is, on this generous accounting, close to noise; the same gap on a five-hundred-task set is not. Real task-to-task difficulty variation only widens these intervals further, so treat the numbers above as a floor on the uncertainty, not the whole of it. The lesson is not to distrust small benchmarks — tau-bench’s airline domain and GAIA’s Level 3 tier are deliberately small because hard, well-verified tasks in those domains are expensive to build — but to stop reading a two-point movement on either of them as a finding.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to A Practitioner's Map of Agent Evaluation Frameworks

See this formula across 1 published context →

Browse the mathematical compendium →