Equation 14 · Forty Million Clicks Through One Uneven Door
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol f
f is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
where is the “data defect correlation” between selection and outcome across the full population, f = n/N is the sampling fraction, and is the outcome’s population standard deviation. The identity’s uncomfortable lesson for any self-selected big-data project is that the error term does not shrink as n grows unless f grows with it — and for a national population sampled through a viral website, f stays vanishingly small no matter how large n gets. Meng introduced this identity using the 2016 US presidential election’s Cooperative Congressional Election Study data as its worked example, and the pattern he documents there is the same “Law of Large Populations” this article is…
Read the full surrounding passage
where is the “data defect correlation” between selection and outcome across the full population, f = n/N is the sampling fraction, and is the outcome’s population standard deviation. The identity’s uncomfortable lesson for any self-selected big-data project is that the error term does not shrink as n grows unless f grows with it — and for a national population sampled through a viral website, f stays vanishingly small no matter how large n gets. Meng introduced this identity using the 2016 US presidential election’s Cooperative Congressional Election Study data as its worked example, and the pattern he documents there is the same “Law of Large Populations” this article is applying here: across US states, the larger a state’s voter population, the further the self-selected sample’s estimated Trump vote share tended to sit from the state’s true result relative to the sample’s own confidence interval — bigger population sizes made the bias worse, not better, because a fixed small defect correlation gets multiplied by a sampling-fraction term that shrinks fastest exactly where the population is largest [ 11 ] . Nothing about that mechanism is specific to elections; it is a property of the arithmetic, and a 39-million-decision survey spread across 130 national populations of wildly different sizes sits squarely inside the class of designs the identity describes.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.