Published equation contexts
Why this formula appears here
That instability is worth making precise rather than gesturing at. If a benchmark of n instances is treated as n independent Bernoulli trials with true resolve probability p , the standard error of the observed proportion is . For n = 500 and near one half, this is on the order of two percentage points — before accounting for the additional variance introduced by sampling temperature, agentic scaffolding, or a different number of retries per problem. A published gap of four or five points between two systems evaluated under different harnesses, different reasoning-effort settings, and different retry budgets is not obviously a capability gap at all; it may be…
Read the representative guide
Symbol n
n occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Read this term in its guide →Numerator: hat p (1 - hat p)
The complete quantity above the fraction bar.
Read this term in its guide →How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
Research cited beside this formula
Published contexts (1)
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 5 · AI Agents & Systems
Measuring Claude Code and Agentic Development Tools: Evidence, Benchmarks, and Uncertainty
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
That instability is worth making precise rather than gesturing at. If a benchmark of n instances is treated as n independent Bernoulli trials with true resolve probability p , the standard error of the observed proportion is . For n = 500 and near one half, this is on the order of two percentage points — before accounting for the additional variance introduced by sampling temperature, agentic scaffolding, or a different number of retries per problem. A published gap of four or five points between two systems evaluated under different harnesses, different reasoning-effort settings, and different retry budgets is not obviously a capability gap at all; it may be…