← Back to article

Equation 20 · Forty Million Clicks Through One Uneven Door

What does this equation mean?

CSDBc(δ)=δσYfc1−fc.\mathrm{CSDB}_c(\delta) = \frac{\delta}{\sigma_Y}\sqrt{\frac{f_c}{1-f_c}} .

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Start withdelta
Divide bysigma_Y
This relates toCSDB_c(delta)
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

cc

Symbol c

c is part of the quantity the equation computes from the expression on the right.

Understand this part →

δ\delta

Symbol delta

the given bias.

Understand this part →

σY\sigma_Y

Symbol sigma_Y

the outcome’s population standard deviation.

Understand this part →

fcf_c

Symbol f_c

fcf_c is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
√

√

Take a square root.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

1−fc1-f_c

Denominator: 1-f_c

The complete quantity below the fraction bar; it must be nonzero for this division.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Solving for the minimum defect correlation needed to produce a given bias δ\delta , and taking the most generous possible bound for a roughly binary preference indicator, σY\sigma_Y ≤\le 0.5 , gives what this article calls, for a specific country c , its Country Selection-Divergence Bound: CSDBc(δ)=δσYfc1−fc\mathrm{CSDB}_c(\delta) = \frac{\delta}{\sigma_Y}\sqrt{\frac{f_c}{1-f_c}} . This is a DERIVED quantity, not a measurement: it has not been run against real per-country population totals or the actual SharedResponsesSurvey.csv microdata, and no value of ρR,Y\rho_{R,Y} for any real country is claimed here. What can be computed honestly, using only the paper’s own reported extremes for n — 101 respondents at the low end, 448,125 at the high end [ 1 ] — is how small…
Read the full surrounding passage
Solving for the minimum defect correlation needed to produce a given bias δ\delta , and taking the most generous possible bound for a roughly binary preference indicator, σY\sigma_Y ≤\le 0.5 , gives what this article calls, for a specific country c , its Country Selection-Divergence Bound: CSDBc(δ)=δσYfc1−fc\mathrm{CSDB}_c(\delta) = \frac{\delta}{\sigma_Y}\sqrt{\frac{f_c}{1-f_c}} . This is a DERIVED quantity, not a measurement: it has not been run against real per-country population totals or the actual SharedResponsesSurvey.csv microdata, and no value of ρR,Y\rho_{R,Y} for any real country is claimed here. What can be computed honestly, using only the paper’s own reported extremes for n — 101 respondents at the low end, 448,125 at the high end [ 1 ] — is how small ρR,Y\rho_{R,Y} would need to be, at illustrative (not dataset-matched) population scales, to produce a cross-country divergence of δ\delta = 0.1 , a magnitude in the range visible between clusters in the paper’s own Fig. 3b radar plot. For a nation of order 10^{6} residents sampled at n=101 : f ≈\approx 1.01×\times10^{-4} , giving CSDB\mathrm{CSDB} ≈\approx 0.002 . For a nation of order 3×\times10^{8} residents sampled at n=448{,}125 : f ≈\approx 1.49×\times10^{-3} , giving CSDB\mathrm{CSDB} ≈\approx 0.0077 . Both bounds sit an order of magnitude below the largest within-sample demographic coefficient the paper itself reports (0.091) as too small to worry about. If a correlation on the order of two-thousandths to eight-thousandths between “reaches the Moral Machine” and “holds this particular moral preference” is enough to move a country’s measured position by as much as the clusters differ, the paper’s own demonstration that observed demographics barely move the estimates does not bound the unobserved selection process nearly as tightly as it appears to. Using a more conservative σY\sigma_Y = 0.3 , reflecting that most individual preference indicators are not split near 50/50, raises both bounds by roughly two-thirds — to about 0.003 and 0.013 — without changing the conclusion that the required leak is tiny.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Forty Million Clicks Through One Uneven Door

See this formula across 1 published context →

Browse the mathematical compendium →