Equation 4 · A Practitioner's Map of Agent Evaluation Frameworks
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol n
n is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
gives a rough sense of how much noise sits under a given n , even granting the generous and almost certainly false assumption that every task in the set is an equally difficult coin flip. At the widest plausible spread, = 0.5 , a 95% interval ( 1.96 ) works out to roughly 13.9 points on a fifty-task set — close to the size of tau-bench’s original airline domain or GAIA’s seventy-five-question Level 3 tier — versus roughly 4.4 points on SWE-bench Verified’s five hundred tasks, or 5.1 points on OSWorld’s 369. A three- or four-point difference between two systems on a fifty- or seventy-five-task slice is, on this generous accounting, close to noise; the…
Read the full surrounding passage
gives a rough sense of how much noise sits under a given n , even granting the generous and almost certainly false assumption that every task in the set is an equally difficult coin flip. At the widest plausible spread, = 0.5 , a 95% interval ( 1.96 ) works out to roughly 13.9 points on a fifty-task set — close to the size of tau-bench’s original airline domain or GAIA’s seventy-five-question Level 3 tier — versus roughly 4.4 points on SWE-bench Verified’s five hundred tasks, or 5.1 points on OSWorld’s 369. A three- or four-point difference between two systems on a fifty- or seventy-five-task slice is, on this generous accounting, close to noise; the same gap on a five-hundred-task set is not. Real task-to-task difficulty variation only widens these intervals further, so treat the numbers above as a floor on the uncertainty, not the whole of it. The lesson is not to distrust small benchmarks — tau-bench’s airline domain and GAIA’s Level 3 tier are deliberately small because hard, well-verified tasks in those domains are expensive to build — but to stop reading a two-point movement on either of them as a finding.
Sources cited in the article section
These citations give research context. Read each source to check which claims it supports.
Return to A Practitioner's Map of Agent Evaluation Frameworks