← Back to article

Equation 5 · A Practitioner's Map of Agent Evaluation Frameworks

What does this equation mean?

p^=0.5\hat p = 0.5

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operations0.5
Result or conditionhat p
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

p^\hat p

Symbol hat p

hat p is part of the quantity the equation computes from the expression on the right.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

gives a rough sense of how much noise sits under a given n , even granting the generous and almost certainly false assumption that every task in the set is an equally difficult coin flip. At the widest plausible spread, p^\hat p = 0.5 , a 95% interval ( zα/2z_{\alpha/2} ≈\approx 1.96 ) works out to roughly ±\pm 13.9 points on a fifty-task set — close to the size of tau-bench’s original airline domain or GAIA’s seventy-five-question Level 3 tier — versus roughly ±\pm 4.4 points on SWE-bench Verified’s five hundred tasks, or ±\pm 5.1 points on OSWorld’s 369. A three- or four-point difference between two systems on a fifty- or seventy-five-task slice is, on this generous accounting, close to noise; the…
Read the full surrounding passage
gives a rough sense of how much noise sits under a given n , even granting the generous and almost certainly false assumption that every task in the set is an equally difficult coin flip. At the widest plausible spread, p^\hat p = 0.5 , a 95% interval ( zα/2z_{\alpha/2} ≈\approx 1.96 ) works out to roughly ±\pm 13.9 points on a fifty-task set — close to the size of tau-bench’s original airline domain or GAIA’s seventy-five-question Level 3 tier — versus roughly ±\pm 4.4 points on SWE-bench Verified’s five hundred tasks, or ±\pm 5.1 points on OSWorld’s 369. A three- or four-point difference between two systems on a fifty- or seventy-five-task slice is, on this generous accounting, close to noise; the same gap on a five-hundred-task set is not. Real task-to-task difficulty variation only widens these intervals further, so treat the numbers above as a floor on the uncertainty, not the whole of it. The lesson is not to distrust small benchmarks — tau-bench’s airline domain and GAIA’s Level 3 tier are deliberately small because hard, well-verified tasks in those domains are expensive to build — but to stop reading a two-point movement on either of them as a finding.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to A Practitioner's Map of Agent Evaluation Frameworks

See this formula across 1 published context →

Browse the mathematical compendium →