Equation 5 · A Practitioner's Map of Agent Evaluation Frameworks
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol hat p
hat p is part of the quantity the equation computes from the expression on the right.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
gives a rough sense of how much noise sits under a given n , even granting the generous and almost certainly false assumption that every task in the set is an equally difficult coin flip. At the widest plausible spread, = 0.5 , a 95% interval ( 1.96 ) works out to roughly 13.9 points on a fifty-task set — close to the size of tau-bench’s original airline domain or GAIA’s seventy-five-question Level 3 tier — versus roughly 4.4 points on SWE-bench Verified’s five hundred tasks, or 5.1 points on OSWorld’s 369. A three- or four-point difference between two systems on a fifty- or seventy-five-task slice is, on this generous accounting, close to noise; the…
Read the full surrounding passage
gives a rough sense of how much noise sits under a given n , even granting the generous and almost certainly false assumption that every task in the set is an equally difficult coin flip. At the widest plausible spread, = 0.5 , a 95% interval ( 1.96 ) works out to roughly 13.9 points on a fifty-task set — close to the size of tau-bench’s original airline domain or GAIA’s seventy-five-question Level 3 tier — versus roughly 4.4 points on SWE-bench Verified’s five hundred tasks, or 5.1 points on OSWorld’s 369. A three- or four-point difference between two systems on a fifty- or seventy-five-task slice is, on this generous accounting, close to noise; the same gap on a five-hundred-task set is not. Real task-to-task difficulty variation only widens these intervals further, so treat the numbers above as a floor on the uncertainty, not the whole of it. The lesson is not to distrust small benchmarks — tau-bench’s airline domain and GAIA’s Level 3 tier are deliberately small because hard, well-verified tasks in those domains are expensive to build — but to stop reading a two-point movement on either of them as a finding.
Sources cited in the article section
These citations give research context. Read each source to check which claims it supports.
Return to A Practitioner's Map of Agent Evaluation Frameworks