Forty Million Clicks Through One Uneven Door
Every one of the Moral Machine’s 39.61 million recorded decisions carries a quiet administrative fact that never appears in the paper’s headline claims: the “country” attached to each decision was never asked, only inferred, in real time, from the IP address of the device that loaded the page. As the paper’s own Methods section states it, “the country from which a user accesses the website is determined through the IP address of their computer or mobile device,” and that inference is what feeds “a vector of moral preferences for each country” and the country-level correlations built on top of it [1]. A visitor in Lagos, Lahore, or Ljubljana becomes a data point for Nigeria, Pakistan, or Slovenia because a packet’s return address said so, not because a census taker did. That is a defensible engineering shortcut — there was no other way to geolocate millions of anonymous browser sessions — but it previews the shape of the larger argument that follows: a country’s slice of this dataset is not a sample of that country’s population. It is a sample of whoever, in that country, happened to own a device that could load a nine-dimensional autonomous-vehicle dilemma game built at the MIT Media Lab, translated into ten languages, and shared widely enough to go viral.
The game is the Moral Machine, and the paper is “The Moral Machine experiment,” published in Nature in November 2018 by Edmond Awad, Sohan Dsouza, Richard Kim, Jonathan Schulz, Joseph Henrich, Azim Shariff, Jean-François Bonnefon, and Iyad Rahwan — at the time, Awad and Rahwan both at the MIT Media Lab, Rahwan also at MIT’s Institute for Data, Systems and Society [1]. Awad now holds a senior lectureship in Economics and the Institute for Data Science and AI at the University of Exeter, with further appointments at Oxford’s Uehiro Institute, the Max Planck Institute for Human Development, and the Alan Turing Institute [13]. Rahwan now directs the Center for Humans and Machines at the Max Planck Institute for Human Development in Berlin [14]. Between them and their six co-authors, the experiment asked visitors to choose, scenario after scenario, which of two groups a self-driving car’s failed brakes should kill, varying nine attributes at once — species, number of characters, age, sex, fitness, social status, legality, relation to the vehicle, and whether to act or stay the course — and used conjoint analysis to read off how heavily each attribute weighed in the aggregate choice. The scenario design itself repays a close look, because the critique that follows depends on knowing exactly what varies and what does not. Each of the platform’s thirteen-scenario sessions presents dilemmas built from a pool of twenty character types — man, woman, pregnant woman, baby in a stroller, elderly man, elderly woman, boy, girl, homeless person, large woman, large man, criminal, male and female executives, male and female athletes, male and female doctors, a dog, and a cat — with three structural attributes (whether the car swerves, whether the imperiled group is pedestrians or passengers, and whether a crossing is legal) randomized in every scenario, and six character attributes each independently randomized across a pair of scenarios per session [1, 4]. That randomization is exactly what conjoint analysis needs to isolate each attribute’s causal weight from the others, and it is a genuinely different, more demanding design than the single-dimension surveys that preceded it. By any honest count this is also one of the largest single behavioral datasets ever assembled on a moral question, and it is worth crediting exactly what that scale bought before asking what it did not.
What Forty Million Choices Actually Prove
Conjoint analysis works by randomizing every attribute of every scenario independently, so that the estimated effect of “sparing the young” cannot be confounded with the effect of “sparing more characters” — the two vary independently by design, and the resulting average marginal component effect (AMCE) has a specific, defensible causal interpretation established well before this paper existed [10]. The Moral Machine’s own robustness checks bear this out at a level of thoroughness that is easy to undersell in summary. Extended Data Fig. 1 confirms the design’s three simplifying assumptions hold: effect estimates stay stable regardless of a scenario’s order within a session, regardless of which side of the screen a profile appeared on, and regardless of whether the theoretical or the actually-collected distribution of attribute combinations is used. Extended Data Fig. 2 confirms the same estimates hold whether or not a respondent read the scenario’s text description, whether they used a desktop or a mobile device, and whether the analysis uses every recorded response or only fully completed thirteen-scenario sessions. None of this is decorative. Each check closes a specific, nameable way the estimates could have been artifacts of the interface rather than the choices.
The individual-level analysis is similarly careful about what it claims. Of the 492,921 Moral Machine users who completed the platform’s optional demographic survey, regression estimates including age, sex, income, education, political orientation, and religiosity as covariates find that none of the nine attribute effects moves by more than roughly a tenth of a probability point for any one demographic split — the single largest of the forty-two reported coefficients is 0.091, on religiosity’s association with sparing humans over pets [1]. That is a genuinely useful, and genuinely reassuring, individual-level finding: whatever is driving the strong, consistent global preferences this platform surfaced — for sparing humans, sparing more lives, sparing the young — it is not primarily a story about which kind of person, among the millions who reached the site, happened to answer the demographic questions a particular way. The paper is not naive about the tool it built. Its Discussion states outright that “our sample is self-selected, and not guaranteed to exactly match the socio-demographics of each country,” and that “policymakers should not embrace our data as the final word on societal preferences — even if our sample is arguably close to the internet-connected, tech-savvy population who is interested in driverless car technology, and more likely to participate in early adoption” [1]. Very little of what follows in this article is an argument the authors would find surprising in kind. It is an argument about how far one particular sentence of theirs can be made to travel, and about a test of it that has a design sitting in the same appendix that names the problem.
The Turnstile Nobody Weighed at the Door
Name the single inferential step the rest of this article is built around, because everything downstream depends on stating it precisely rather than gesturing at “selection bias” in general. The paper reports three cross-country “moral clusters” — Western, Eastern, and Southern — built by hierarchically clustering 130 countries’ nine-dimensional AMCE vectors, each country represented by at least 100 respondents, ranging up to a reported maximum of 448,125 [1]. It further reports that a country’s position in this space correlates with its individualism score, its rule-of-law score, and its economic inequality: an ordinary least-squares regression across the 56 countries with full covariate data finds individualism predicting the preference for sparing more characters at
For that treatment to be licensed, one assumption has to hold: that the process selecting who, within each country, ends up as a Moral Machine respondent distorts each country’s underlying nine-dimensional preference vector by roughly the same amount and in roughly the same direction, so that the differences between countries’ sample vectors still track the differences between countries’ population vectors. Call the alternative the differential-turnstile model. Every visitor who reaches the Moral Machine passes through the same conceptual turnstile — an internet connection, enough curiosity about the premise to click through, sufficient fluency in one of ten offered languages, enough free attention to complete thirteen forced-choice dilemmas — but the rate at which each country’s population is admitted through that turnstile is not remotely uniform. The platform’s own aggregate respondent profile, reported in Extended Data Fig. 6, is roughly 70% male against 28% female and 2% “other,” with an age distribution peaking somewhere in the late teens and continuing to fall through the twenties and thirties, and with more than a third of respondents reporting the platform’s lowest income bracket — a profile the paper’s own caption describes candidly as showing users who “went through college, and are in their 20s or 30s” [1]. That is a demographic stratum that international survey research consistently finds closer, across countries, to a shared “young, educated, online” pole than the general populations it is drawn from [6, 7]. If the size of that stratum, relative to each country’s total population, itself varies with GDP, internet penetration, and institutional development — the same macro variables the paper correlates with its moral clusters — then a purely compositional artifact of who got through the door would reproduce the reported pattern without requiring any difference in the underlying population’s moral preferences at all.
One Country Checked, One Hundred Thirty Assumed
The paper does not ignore this concern; it tests a version of it, and the test is worth describing exactly rather than dismissing. Extended Data Fig. 7 compares the AMCEs computed from the raw US sample against AMCEs recomputed after post-stratifying that same sample to match the age, sex, income, and education margins of the American Community Survey. “Except for ‘Relation to AV,’” the caption reports, “the direction and order of all effects are unaffected” [1]. That is a real, useful robustness result, and it demonstrates something specific: within the United States, the demographic skew of who reached the Moral Machine — visibly more male, younger, and more college-educated than the ACS baseline across every one of the four panels in Extended Data Fig. 7a–d — does not appear to have distorted the ordering of the nine preferences enough to matter. Only one effect, “Relation to AV,” changed at all, and it was already the second-smallest effect in the whole ranking before reweighting.
What that result cannot demonstrate is the thing the cross-country claims actually require. Post-stratification corrects for imbalance on covariates that are observed and matched between sample and target population — here, four demographic margins, for one country. It is silent on any bias correlated with variables outside that adjustment set, and the survey methods literature on this exact procedure is explicit that its guarantee is only as good as the ignorability assumption it rests on: unless the propensity to respond is fully explained by the variables being balanced, weighting on those variables leaves whatever remains untouched [12]. The Moral Machine’s country-level claims do not depend on within-country demographic imbalance behaving consistently — they depend on a claim never checked in any country at all: that the fraction of each nation’s population thin enough, wired enough, and moved enough to complete the exercise is unrelated to GDP, rule of law, and inequality, or if related, distorts moral preferences identically regardless of country. A card catalogue of one country’s census numbers, checked once, says nothing about the other hundred and twenty-nine nations whose census books were never opened for this purpose at all.
The Reliability Check That Doubles as the Finding
The paper’s Discussion contains a sentence that, read closely, states the circularity this article is built around more precisely than any external critique needs to: “the fact that the cross-societal variation we observed aligns with previously established cultural clusters, as well as the fact that macro-economic variables are predictive of Moral Machine responses, are good signs about the reliability of our data, as is our post-stratification” [1]. Read that sentence twice. The same two correlations — clusters aligning with known cultural groupings, and macro-economic variables predicting the moral-preference vectors — are offered in one breath as evidence that the data collection process can be trusted, and in the next breath (via Fig. 4 and Extended Data Table 2) as the headline empirical result the paper is reporting to the world. A quantity cannot serve as both the calibration check on an instrument and the reading that instrument produces. If national wealth and institutional quality predict who reaches the Moral Machine website in the first place — which digital-divide statistics on internet penetration make close to certain in some degree — then national wealth and institutional quality predicting the sample’s moral-preference vector is exactly what a pure selection artifact would look like, indistinguishable in this evidence from a real population-level moral difference. The alignment with the Inglehart–Welzel cultural map is genuinely striking and genuinely worth reporting. It is just as consistent with “richer, more institutionally developed countries send a more similar-looking slice of their population through the turnstile” as with “richer, more institutionally developed countries hold more similar moral views.”
This is also where the paper’s own hierarchical clustering validation becomes double-edged. Extended Data Fig. 5 reports the Ward clustering’s purity against 130 countries’ true (arbitrary) labels at 0.5888, and its maximum bipartite matching against random ninefold assignment at 0.5421 — both comfortably above the null distributions the authors compute by randomly reassigning countries to clusters [1]. That comparison correctly establishes the clusters are not statistical noise. It says nothing about whether the clusters are moral or demographic in origin, because a differential-turnstile process would produce non-random, above-chance clustering too — clusters that track which countries’ internet-connected, younger populations happen to resemble each other, which is exactly the same non-random structure a real cross-cultural moral difference would produce. The silhouette index in the same figure sits mostly in the 0.10–0.20 range across cluster counts — real structure, but weak structure, of a magnitude compatible with either account.
It is also worth noting, in fairness to the paper, that Extended Data Table 2’s four regressions do not read as a story fitted to confirm itself: several of the sixteen reported coefficients are small and statistically indistinguishable from zero, including individualism’s association with sparing higher-status characters (
The Arithmetic of How Small a Leak Would Have to Be
A critique earns the right to be taken seriously only once it can say how large the effect it is worried about would need to be, and Xiao-Li Meng’s 2018 identity for bias in self-selected big-data samples gives a way to say exactly that, using only numbers the paper and its data-availability statement already make public [11]. For a population of size
where
Solving for the minimum defect correlation needed to produce a given bias
This is a DERIVED quantity, not a measurement: it has not been run against real per-country population totals or the actual SharedResponsesSurvey.csv microdata, and no value of
What Would Have to Be True for the Clusters to Survive
State the strongest form of the objection this article is going to lose to, if it loses. The paper’s own evidence already suggests the turnstile model may not have much room to work with: within-sample demographic coefficients cap out at 0.091, and the one post-stratification actually run changed almost nothing. If demographics this coarse barely move the estimates inside a country, why would a subtler selection process move them between countries? The honest answer is that these are different questions, and selection-induced range restriction predicts exactly this pattern regardless of which account is right. A platform reached only by a narrow global stratum — young, educated, online — will show little internal variation on stratum-level traits like income or education, because everyone who got through the turnstile already resembles everyone else who did on those traits. That tells us nothing about whether the width of the door itself — the fraction of each country’s population thin enough to fit through it — correlates with GDP, institutions, and inequality in a way that shapes the cluster structure. A single country’s within-sample check cannot speak to a between-country selection process it was never designed to detect.
The decisive test follows directly, uses no data the platform’s own availability statement does not already promise, and has — after adversarial search — not been run by these authors or by anyone else. Post-stratify each of the 130 countries’ Moral Machine samples to national census age-by-sex-by-education margins, the only three demographic fields collected by the platform’s survey that have a genuine national-census analogue [4] — political orientation and religiosity, both collected as 0-to-1 sliders, have no comparable national benchmark and cannot be post-stratified against anything. Recompute each country’s nine AMCEs on the reweighted sample, rerun the Ward hierarchical clustering, and rerun the Extended Data Table 2 regressions of individualism, rule of law, and inequality against the reweighted country vectors. Alongside this, contrast respondents who completed all thirteen scenarios against those who dropped out partway, within each country, as a direct probe of whether the kind of person who finishes differs from the kind who does not in a way that itself tracks national development. If the three-cluster partition recurs with an adjusted Rand index at or above 0.8 against the published clusters, if the individualism, rule-of-law, and inequality associations keep their sign and remain distinguishable from zero under cluster-robust standard errors, and if the completer/non-completer contrasts stay small relative to the distances separating the clusters, this article’s central claim fails and the cross-country findings should be read as exactly what they are billed as. If the clusters or the macro-correlations move substantially under reweighting, the opposite follows. This article has run neither half of that test. It specifies what would settle the question using data the paper’s own authors have already said, in print, “can be used beyond replication to answer follow-up research questions” [1]. None of the needed material sits behind a request form or an embargo: the OSF repository hosts the full individual-level response file, the demographic-survey file merged to it by session, and the country-cluster and effect-size files the paper’s own figures were built from, openly and without access restriction [4]. The demographic-survey file alone documents, field by field, exactly which self-reported variables — age (collected as free-text and limited to the 18-to-75 range in that file), gender, education, income, political orientation, and religiosity — are available to condition any reweighting on, and just as usefully documents which are not: political orientation and religiosity are continuous sliders defaulting to the scale’s midpoint when unanswered, with no equivalent category in any national census this article is aware of, which is why the reweighting proposed here is explicitly restricted to the three fields — age, sex, and education — that a national statistics office actually publishes [4]. A second, smaller measurement problem sits underneath even that restricted reweighting and is worth naming rather than quietly assuming away: the “country” being reweighted is itself an IP-geolocation artifact, not a self-report, so a respondent traveling, tunneling through a foreign server, or using a mobile carrier’s out-of-country routing would be silently reassigned to the wrong national population before the reweighting ever begins. That measurement noise likely adds a small amount of unrelated error in both directions rather than a systematic bias favoring the turnstile account over the population account, but a careful reanalysis should quantify it rather than assume it away.
What the Turnstile Model Does Not Explain Away
None of this touches the paper’s global, within-sample findings, and a critique that let readers forget that would be doing them a disservice. The preference for sparing humans over pets, for sparing more lives, and for sparing the young are not artifacts of cross-country comparison — they are aggregate preferences within the platform’s own population, and they triangulate independently. The authors’ own Reply to a later comment reports that 585,531 users who separately positioned continuous preference sliders — a method built specifically to let respondents express an equal-treatment preference the forced-choice format could not — reproduced the same ranking, confirming the four strongest scenario-based preferences as strong and the four weakest as weak [3]. That comment, from Yochanan Bigman and Kurt Gray, had shown that when offered an explicit equality option, 97.9% of respondents preferred to treat men and women equally, against the 87.7% who chose to save women when the forced-choice format gave them no equality option at all [2] — a real and important limitation of the forced-choice paradigm’s ability to detect egalitarian preferences, one the original authors accepted and engaged rather than dismissed, while pointing out in their Reply that framing effects cut in both directions in Bigman and Gray’s own studies as well [3]. None of this bears on the cross-country clustering claim this article is about, but all of it is evidence of a research program that keeps testing its own instrument rather than resting on one result.
That pattern continues. Joseph Henrich, the Moral Machine’s fifth co-author, is also the lead author of the paper that gave the acronym WEIRD — Western, Educated, Industrialized, Rich, Democratic — to exactly the kind of unrepresentative convenience sample this article argues the Moral Machine’s country-level comparisons may still be built on, and a co-author, two years after the Moral Machine paper itself, of a follow-up that built an actual cultural-and-psychological-distance metric from the same concern [6, 7]. That is not a contradiction to expose; it is worth stating plainly, without irony, as evidence that the people who built this platform have spent their careers on the exact statistical hazard this article raises against their own dataset. Awad’s own 2025 methods essay makes the point even more directly about the Moral Machine specifically, describing its own platform’s central limitation in language this article did not need to improve on: the approach “has problems with representation (internet users, highly educated),” and “highly educated participants from non-Western countries show similarity to Western participants… and so relying on such participants can understate or otherwise fail to represent cross-cultural differences” [5, 8]. That sentence, written by the platform’s own lead architect seven years after publication, is the differential-turnstile model stated in the first person.
The genuine tension the paper surfaces between its own strongest, most consistent finding — a broad preference for sparing the young — and German Ethical Rule 9’s prohibition on any distinction based on personal features also survives this critique untouched [1]. That tension is about whether a real, robustly measured public preference should ever be encoded into policy, not about whether the preference’s cross-country distribution is measured correctly. Both questions matter. They are not the same question, and only one of them is this article’s target.
The Census Nobody Has Opened Yet
A wall map dense with pins in some places and bare in others is not, by itself, evidence about the places with few pins. It is evidence about where the map’s own postal service reached. The Moral Machine’s country-level clusters and their macro-economic correlations may well be exactly what they are presented as — a genuine, if partial, cartography of how national populations differ in the moral weight they place on age, law, and status. They may also be, in whole or in part, a cartography of where a research team’s viral experiment found an audience thin enough to fit through its own turnstile at a rate that happens to track GDP, rule of law, and inequality for reasons that have nothing to do with morality at all. Both readings fit every number this article has quoted. The paper’s own released microdata, its own country-level sample counts, and its own demographic-survey fields are sufficient to tell the two readings apart; nobody, including the team that built the platform, has yet run that comparison. The turnstile has been counting for eight years. The census that would tell us what it missed is sitting, unopened, in the same public dataset the original authors have never stopped inviting other researchers to use.