← All parts of this equation

Equation 10 · Part 5 · Measuring Frontier Models: Contamination, Variance, and What a Score Can Support

√

SE=p(1−p)n,\mathrm{SE} = \sqrt{\frac{p(1-p)}{n}} ,
√

What this part means

Take a square root.

Its job in the formula

Take a square root.

The passage around this formula

For a suite of n independent items with true rate p , the standard error of the estimate is SE=p(1−p)n\mathrm{SE} = \sqrt{\frac{p(1-p)}{n}} . which for p ≈\approx 0.5 and n = 200 is about 3.5 percentage points — before adding any run-to-run generation variance, and before accounting for the fact that benchmark items are not independent. A reported two-point difference between two systems on such a suite is not evidence of anything.

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.