← Back to article

Equation 9 · Forty Million Clicks Through One Uneven Door

What does this equation mean?

Yˉn−YˉN=ρR,Y1−ff σY,\bar{Y}_n - \bar{Y}_N = \rho_{R,Y}\sqrt{\frac{1-f}{f}}\,\sigma_Y ,

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Start with1-f
Divide byf
This relates tobarY_n - barY_N
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

Yˉn\bar{Y}_n

Symbol barY_n

barYnY_n is part of the quantity the equation computes from the expression on the right.

Understand this part →

YˉN\bar{Y}_N

Symbol barY_N

barYNY_N is part of the quantity the equation computes from the expression on the right.

Understand this part →

ρR,Y\rho_{R,Y}

Symbol rho_R,Y

the no value of.

Understand this part →

ff

Symbol f

f is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

σY\sigma_Y

Symbol sigma_Y

the outcome’s population standard deviation.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
√

√

Take a square root.

Understand this part →

subtraction

subtraction

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

1−f1-f

Numerator: 1-f

The complete quantity above the fraction bar.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

A critique earns the right to be taken seriously only once it can say how large the effect it is worried about would need to be, and Xiao-Li Meng’s 2018 identity for bias in self-selected big-data samples gives a way to say exactly that, using only numbers the paper and its data-availability statement already make public [ 11 ] . For a population of size N with a binary response indicator R (did this person’s data reach the sample) and an outcome Y (their realized value on some Moral-Machine-relevant preference indicator), Meng’s identity relates the sample mean to the true population mean as Yˉn−YˉN=ρR,Y1−ff σY\bar{Y}_n - \bar{Y}_N = \rho_{R,Y}\sqrt{\frac{1-f}{f}}\,\sigma_Y . where ρR,Y\rho_{R,Y} is the “data defect correlation” between selection and outcome…
Read the full surrounding passage
A critique earns the right to be taken seriously only once it can say how large the effect it is worried about would need to be, and Xiao-Li Meng’s 2018 identity for bias in self-selected big-data samples gives a way to say exactly that, using only numbers the paper and its data-availability statement already make public [ 11 ] . For a population of size N with a binary response indicator R (did this person’s data reach the sample) and an outcome Y (their realized value on some Moral-Machine-relevant preference indicator), Meng’s identity relates the sample mean to the true population mean as Yˉn−YˉN=ρR,Y1−ff σY\bar{Y}_n - \bar{Y}_N = \rho_{R,Y}\sqrt{\frac{1-f}{f}}\,\sigma_Y . where ρR,Y\rho_{R,Y} is the “data defect correlation” between selection and outcome across the full population, f = n/N is the sampling fraction, and σY\sigma_Y is the outcome’s population standard deviation. The identity’s uncomfortable lesson for any self-selected big-data project is that the error term does not shrink as n grows unless f grows with it — and for a national population sampled through a viral website, f stays vanishingly small no matter how large n gets. Meng introduced this identity using the 2016 US presidential election’s Cooperative Congressional Election Study data as its worked example, and the pattern he documents there is the same “Law of Large Populations” this article is applying here: across US states, the larger a state’s voter population, the further the self-selected sample’s estimated Trump vote share tended to sit from the state’s true result relative to the sample’s own confidence interval — bigger population sizes made the bias worse, not better, because a fixed small defect correlation gets multiplied by a sampling-fraction term that shrinks fastest exactly where the population is largest [ 11 ] . Nothing about that mechanism is specific to elections; it is a property of the arithmetic, and a 39-million-decision survey spread across 130 national populations of wildly different sizes sits squarely inside the class of designs the identity describes.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Forty Million Clicks Through One Uneven Door

See this formula across 1 published context →

Browse the mathematical compendium →