← Back to article

Equation 11 · Forty Million Clicks Through One Uneven Door

What does this equation mean?

f=n/Nf = n/N

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationsn/N
Result or conditionf
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

ff

Symbol f

f is part of the quantity the equation computes from the expression on the right.

Understand this part →

nn

Symbol n

n is an input to the expression that computes the quantity on the left.

Understand this part →

NN

Symbol N

the sampling fraction.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

where ρR,Y\rho_{R,Y} is the “data defect correlation” between selection and outcome across the full population, f = n/N is the sampling fraction, and σY\sigma_Y is the outcome’s population standard deviation. The identity’s uncomfortable lesson for any self-selected big-data project is that the error term does not shrink as n grows unless f grows with it — and for a national population sampled through a viral website, f stays vanishingly small no matter how large n gets. Meng introduced this identity using the 2016 US presidential election’s Cooperative Congressional Election Study data as its worked example, and the pattern he documents there is the same “Law of Large Populations” this article is…
Read the full surrounding passage
where ρR,Y\rho_{R,Y} is the “data defect correlation” between selection and outcome across the full population, f = n/N is the sampling fraction, and σY\sigma_Y is the outcome’s population standard deviation. The identity’s uncomfortable lesson for any self-selected big-data project is that the error term does not shrink as n grows unless f grows with it — and for a national population sampled through a viral website, f stays vanishingly small no matter how large n gets. Meng introduced this identity using the 2016 US presidential election’s Cooperative Congressional Election Study data as its worked example, and the pattern he documents there is the same “Law of Large Populations” this article is applying here: across US states, the larger a state’s voter population, the further the self-selected sample’s estimated Trump vote share tended to sit from the state’s true result relative to the sample’s own confidence interval — bigger population sizes made the bias worse, not better, because a fixed small defect correlation gets multiplied by a sampling-fraction term that shrinks fastest exactly where the population is largest [ 11 ] . Nothing about that mechanism is specific to elections; it is a property of the arithmetic, and a 39-million-decision survey spread across 130 national populations of wildly different sizes sits squarely inside the class of designs the identity describes.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Forty Million Clicks Through One Uneven Door

See this formula across 1 published context →

Browse the mathematical compendium →