← All parts of this equation

Equation 5 · Part 7 · Measuring Claude Code and Agentic Development Tools: Evidence, Benchmarks, and Uncertainty

Numerator: hat p (1 - hat p)

SE(p^)=p^(1−p^)n.\mathrm{SE}(\hat p) = \sqrt{\frac{\hat p (1 - \hat p)}{n}}.
p^(1−p^)\hat p (1 - \hat p)

What this part means

The complete quantity above the fraction bar.

Its job in the formula

hat p (1 - hat p) occurs above the fraction bar. The numerator is divided by the entire denominator below it.

The passage around this formula

That instability is worth making precise rather than gesturing at. If a benchmark of n instances is treated as n independent Bernoulli trials with true resolve probability p , the standard error of the observed proportion p^\hat p is SE(p^)=p^(1−p^)n\mathrm{SE}(\hat p) = \sqrt{\frac{\hat p (1 - \hat p)}{n}}. For n = 500 and p^\hat p near one half, this is on the order of two percentage points — before accounting for the additional variance introduced by sampling temperature, agentic scaffolding, or a different number of retries per problem. A published gap of four or five points between two systems evaluated under different harnesses, different reasoning-effort settings, and different retry budgets is not obviously a capability gap at all; it may be…

Read this part in the article →

Learn the underlying idea

A fraction a/b means a divided by b. The top number is the numerator; the bottom number is the denominator, and it cannot be zero.

Open the illustrated fractions: division written vertically guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.