← Mathematical compendium

Published equation contexts

SEcluster=SEnaive×1+(m−1)ρ\mathrm{SE}_{\text{cluster}} = \mathrm{SE}_{\text{naive}} \times \sqrt{1 + (m - 1)\rho}

Why this formula appears here

5. Non-independence between repeated trials. The confidence-interval formula above assumes each trial is independent of the others, and that assumption is often false in ways that specifically inflate confidence. Miller’s paper measures this directly: several widely used evaluation sets draw multiple questions from a shared context — several questions about the same passage, several sub-tasks from the same underlying scenario — and properly accounting for that clustering, rather than treating every question as its own independent draw, produced clustered standard errors “over 3x larger than naive standard errors” on the affected evals [ 8 ] . The general relationship is the one long used in…

Read the full article-specific guide →

Read the representative guide

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

SEcluster=SEnaive×1+(m−1)ρ.\mathrm{SE}_{\text{cluster}} = \mathrm{SE}_{\text{naive}} \times \sqrt{1 + (m - 1)\rho}.

Equation 18 · Model Evaluation

Ten Ways an Agent Evaluation Can Mislead You Even When It's Working Correctly

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

5. Non-independence between repeated trials. The confidence-interval formula above assumes each trial is independent of the others, and that assumption is often false in ways that specifically inflate confidence. Miller’s paper measures this directly: several widely used evaluation sets draw multiple questions from a shared context — several questions about the same passage, several sub-tasks from the same underlying scenario — and properly accounting for that clustering, rather than treating every question as its own independent draw, produced clustered standard errors “over 3x larger than naive standard errors” on the affected evals [ 8 ] . The general relationship is the one long used in…

Equation guide → · Article →