Equation 10 · Measuring Frontier Models: Contamination, Variance, and What a Score Can Support
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol p
p occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Symbol n
n occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
For a suite of n independent items with true rate p , the standard error of the estimate is . which for p 0.5 and n = 200 is about 3.5 percentage points — before adding any run-to-run generation variance, and before accounting for the fact that benchmark items are not independent. A reported two-point difference between two systems on such a suite is not evidence of anything.
Sources cited in the article section
- [10] Model guidance ↗
- [8] DeepSeek-R1 Incentivizes Reasoning in LLMs through Reinforcement Learning ↗
These citations give research context. Read each source to check which claims it supports.
Return to Measuring Frontier Models: Contamination, Variance, and What a Score Can Support