← All parts of this equation

Equation 5 · Part 1 · Measuring Claude Code and Agentic Development Tools: Evidence, Benchmarks, and Uncertainty

Symbol hat p

SE(p^)=p^(1−p^)n.\mathrm{SE}(\hat p) = \sqrt{\frac{\hat p (1 - \hat p)}{n}}.
p^\hat p

What this part means

the standard error of the observed proportion.

Its job in the formula

hat p occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Where the article explains it

If a benchmark of n instances is treated as n independent Bernoulli trials with true resolve probability p , the standard error of the observed proportion p^\hat p is SE(p^)=p^(1−p^)n\mathrm{SE}(\hat p) = \sqrt{\frac{\hat p (1 - \hat p)}{n}}.

The passage around this formula

…instability is worth making precise rather than gesturing at. If a benchmark of n instances is treated as n independent Bernoulli trials with true resolve probability p , the standard error of the observed proportion p^\hat p is SE(p^)=p^(1−p^)n\mathrm{SE}(\hat p) = \sqrt{\frac{\hat p (1 - \hat p)}{n}}. For n = 500 and p^\hat p near one half, this is on the order of two percentage points — before accounting for the additional variance introduced by sampling temperature, agentic scaffolding, or a different number of retries per problem. A published gap of four or five…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.