← Back to article

Equation 6 · A Practitioner's Map of Agent Evaluation Frameworks

What does this equation mean?

zα/2≈1.96z_{\alpha/2} \approx 1.96

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

zα/2z_{\alpha/2}

Symbol z_α/2

z_α/2 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

≈

≈

Approximately equal to; the equality is not exact.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Its accuracy depends on the assumptions and range of use described in the article.

What the article says around this equation

gives a rough sense of how much noise sits under a given n , even granting the generous and almost certainly false assumption that every task in the set is an equally difficult coin flip. At the widest plausible spread, p^\hat p = 0.5 , a 95% interval ( zα/2z_{\alpha/2} ≈\approx 1.96 ) works out to roughly ±\pm 13.9 points on a fifty-task set — close to the size of tau-bench’s original airline domain or GAIA’s seventy-five-question Level 3 tier — versus roughly ±\pm 4.4 points on SWE-bench Verified’s five hundred tasks, or ±\pm 5.1 points on OSWorld’s 369. A three- or four-point difference between two systems on a fifty- or seventy-five-task slice is, on this generous accounting, close to noise; the…
Read the full surrounding passage
gives a rough sense of how much noise sits under a given n , even granting the generous and almost certainly false assumption that every task in the set is an equally difficult coin flip. At the widest plausible spread, p^\hat p = 0.5 , a 95% interval ( zα/2z_{\alpha/2} ≈\approx 1.96 ) works out to roughly ±\pm 13.9 points on a fifty-task set — close to the size of tau-bench’s original airline domain or GAIA’s seventy-five-question Level 3 tier — versus roughly ±\pm 4.4 points on SWE-bench Verified’s five hundred tasks, or ±\pm 5.1 points on OSWorld’s 369. A three- or four-point difference between two systems on a fifty- or seventy-five-task slice is, on this generous accounting, close to noise; the same gap on a five-hundred-task set is not. Real task-to-task difficulty variation only widens these intervals further, so treat the numbers above as a floor on the uncertainty, not the whole of it. The lesson is not to distrust small benchmarks — tau-bench’s airline domain and GAIA’s Level 3 tier are deliberately small because hard, well-verified tasks in those domains are expensive to build — but to stop reading a two-point movement on either of them as a finding.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to A Practitioner's Map of Agent Evaluation Frameworks

See this formula across 2 published contexts →

Browse the mathematical compendium →