← Back to article

Equation 13 · Forty Million Clicks Through One Uneven Door

What does this equation mean?

nn

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

nn

Symbol n

n is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

where ρR,Y\rho_{R,Y} is the “data defect correlation” between selection and outcome across the full population, f = n/N is the sampling fraction, and σY\sigma_Y is the outcome’s population standard deviation. The identity’s uncomfortable lesson for any self-selected big-data project is that the error term does not shrink as n grows unless f grows with it — and for a national population sampled through a viral website, f stays vanishingly small no matter how large n gets. Meng introduced this identity using the 2016 US presidential election’s Cooperative Congressional Election Study data as its worked example, and the pattern he documents there is the same “Law of Large Populations” this article is…
Read the full surrounding passage
where ρR,Y\rho_{R,Y} is the “data defect correlation” between selection and outcome across the full population, f = n/N is the sampling fraction, and σY\sigma_Y is the outcome’s population standard deviation. The identity’s uncomfortable lesson for any self-selected big-data project is that the error term does not shrink as n grows unless f grows with it — and for a national population sampled through a viral website, f stays vanishingly small no matter how large n gets. Meng introduced this identity using the 2016 US presidential election’s Cooperative Congressional Election Study data as its worked example, and the pattern he documents there is the same “Law of Large Populations” this article is applying here: across US states, the larger a state’s voter population, the further the self-selected sample’s estimated Trump vote share tended to sit from the state’s true result relative to the sample’s own confidence interval — bigger population sizes made the bias worse, not better, because a fixed small defect correlation gets multiplied by a sampling-fraction term that shrinks fastest exactly where the population is largest [ 11 ] . Nothing about that mechanism is specific to elections; it is a property of the arithmetic, and a 39-million-decision survey spread across 130 national populations of wildly different sizes sits squarely inside the class of designs the identity describes.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Forty Million Clicks Through One Uneven Door

Browse the mathematical compendium →