Symbol hat p
hat p occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Read this term in its guide →Published equation contexts
Reported variance, not a bare point estimate. A score computed from n independent trials with an underlying success probability carries a standard error of roughly . and two point estimates whose intervals overlap should not be reported as a ranking. Evan Miller’s statistical treatment of language-model evaluation makes this argument in more general form, framing individual evaluation questions as draws from an unseen larger population and deriving the formulas needed to report genuine uncertainty rather than a single noisy number [ 1 ] . Epoch AI’s benchmarking hub applies the same discipline operationally: it runs each model sixteen times on GPQA Diamond and Mock…
hat p occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Read this term in its guide →n occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Read this term in its guide →The complete quantity above the fraction bar.
Read this term in its guide →With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 14 · Model Evaluation
This equation gives an approximation: it relates the quantities while allowing an approximation.
Reported variance, not a bare point estimate. A score computed from n independent trials with an underlying success probability carries a standard error of roughly . and two point estimates whose intervals overlap should not be reported as a ranking. Evan Miller’s statistical treatment of language-model evaluation makes this argument in more general form, framing individual evaluation questions as draws from an unseen larger population and deriving the formulas needed to report genuine uncertainty rather than a single noisy number [ 1 ] . Epoch AI’s benchmarking hub applies the same discipline operationally: it runs each model sixteen times on GPQA Diamond and Mock…
Equation guide → · Article →