← Mathematical compendium

Published equation contexts

p^  ±  zα/2p^(1−p^)n\hat p \; \pm \; z_{\alpha/2}\sqrt{\frac{\hat p (1-\hat p)}{n}}

Why this formula appears here

The second argument is about how much any one number can actually tell you, independent of harness effects, purely from how many tasks produced it. Treat a benchmark’s headline pass rate as an estimate p^\hat p of a task-set success probability, drawn from n roughly independent trials. The familiar normal approximation to a binomial confidence interval, p^  ±  zα/2p^(1−p^)n\hat p \; \pm \; z_{\alpha/2}\sqrt{\frac{\hat p (1-\hat p)}{n}}. gives a rough sense of how much noise sits under a given n , even granting the generous and almost certainly false assumption that every task in the set is an equally difficult coin flip. At the widest plausible spread, p^\hat p = 0.5 , a 95% interval ( zα/2z_{\alpha/2} ≈\approx 1.96 ) works out to roughly ±\pm 13.9 points on a…

Read the full article-specific guide →

Read the representative guide

zα/2z_{\alpha/2}

Symbol z_α/2

z_α/2 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
nn

Symbol n

n occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

p^  ±  zα/2p^(1−p^)n,\hat p \; \pm \; z_{\alpha/2}\sqrt{\frac{\hat p (1-\hat p)}{n}},

Equation 3 · Model Evaluation

A Practitioner's Map of Agent Evaluation Frameworks

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

The second argument is about how much any one number can actually tell you, independent of harness effects, purely from how many tasks produced it. Treat a benchmark’s headline pass rate as an estimate p^\hat p of a task-set success probability, drawn from n roughly independent trials. The familiar normal approximation to a binomial confidence interval, p^  ±  zα/2p^(1−p^)n\hat p \; \pm \; z_{\alpha/2}\sqrt{\frac{\hat p (1-\hat p)}{n}}. gives a rough sense of how much noise sits under a given n , even granting the generous and almost certainly false assumption that every task in the set is an equally difficult coin flip. At the widest plausible spread, p^\hat p = 0.5 , a 95% interval ( zα/2z_{\alpha/2} ≈\approx 1.96 ) works out to roughly ±\pm 13.9 points on a…

Equation guide → · Article →