← Mathematical compendium

Published equation contexts

SE(p^)≈p^(1−p^)n\mathrm{SE}(\hat p) \approx \sqrt{\frac{\hat p (1-\hat p)}{n}}

Why this formula appears here

Reported variance, not a bare point estimate. A score computed from n independent trials with an underlying success probability p^\hat p carries a standard error of roughly SE(p^)≈p^(1−p^)n\mathrm{SE}(\hat p) \approx \sqrt{\frac{\hat p (1-\hat p)}{n}}. and two point estimates whose intervals overlap should not be reported as a ranking. Evan Miller’s statistical treatment of language-model evaluation makes this argument in more general form, framing individual evaluation questions as draws from an unseen larger population and deriving the formulas needed to report genuine uncertainty rather than a single noisy number [ 1 ] . Epoch AI’s benchmarking hub applies the same discipline operationally: it runs each model sixteen times on GPQA Diamond and Mock…

Read the full article-specific guide →

Read the representative guide

nn

Symbol n

n occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

SE(p^)≈p^(1−p^)n,\mathrm{SE}(\hat p) \approx \sqrt{\frac{\hat p (1-\hat p)}{n}},

Equation 14 · Model Evaluation

How to Actually Compare Frontier AI Models Without Building a Misleading Leaderboard

This equation gives an approximation: it relates the quantities while allowing an approximation.

Reported variance, not a bare point estimate. A score computed from n independent trials with an underlying success probability p^\hat p carries a standard error of roughly SE(p^)≈p^(1−p^)n\mathrm{SE}(\hat p) \approx \sqrt{\frac{\hat p (1-\hat p)}{n}}. and two point estimates whose intervals overlap should not be reported as a ranking. Evan Miller’s statistical treatment of language-model evaluation makes this argument in more general form, framing individual evaluation questions as draws from an unseen larger population and deriving the formulas needed to report genuine uncertainty rather than a single noisy number [ 1 ] . Epoch AI’s benchmarking hub applies the same discipline operationally: it runs each model sixteen times on GPQA Diamond and Mock…

Equation guide → · Article →