Selection tracks investment, not sex, and the theory says so explicitly

Robert Trivers’s 1972 chapter on parental investment and sexual selection is cited by evolutionary psychologists more often than it is read, and the compression costs something. Trivers defined parental investment as “any investment by the parent in an individual offspring that increases the offspring’s chances of surviving (and hence reproducing) at the cost of the parent’s ability to invest in other offspring,” and built an argument that is symmetrical in a way its later popularizations are not: whichever sex invests more in offspring becomes the more discriminating, more competed-for sex, and whichever sex invests less becomes the more competitive, less choosy one [1]. That is a claim about relative investment, not a claim about which sex carries which chromosome.

Trivers’s own examples make the point unmissable. He singled out species with reversed investment patterns — the Phalaropidae, several polyandrous birds, the pipefish-and-seahorse family Syngnathidae, at least one dendrobatid frog — in which the male alone broods the eggs or carries the young, and reported that these species “are striking in showing very high male parental investment correlating with strong sex role reversal: females tend to be more brightly colored, more aggressive and larger than the males, and tend to court them and fight over them” [1]. If the theory predicted “males compete, females choose” as a fact about maleness and femaleness rather than about investment, these species would falsify it outright. Instead they are Trivers’s own supporting evidence, cited specifically to show that the mechanism tracks investment wherever investment happens to fall.

Applied to humans, the asymmetry runs the ordinary way because internal fertilization, a roughly nine-month gestation, and lactation put a floor under female parental investment that has no male equivalent. David Buss and David Schmitt made this the explicit starting premise of Sexual Strategies Theory two decades later, quoting Trivers’s definition directly and noting that women tend to be the more heavily investing sex “in part because fertilization, gestation, and placentation are internal within women,” while a single copulation can cost a man almost nothing [2]. But Buss and Schmitt’s 1993 theory is also routinely flattened into a claim it does not make. The caricature holds that evolutionary psychology predicts blanket promiscuity for men and blanket monogamy-seeking for women; the actual theory proposes that both sexes evolved distinct psychological mechanisms for both short-term and long-term mating, each activated by context, and it devotes entire sections to problems women solve in short-term mating — immediate resource extraction, using brief contact to assess a prospective long-term mate — and to problems men solve in long-term mating, including identifying a partner who shows good parenting skill and genuine willingness to commit [2]. Nine key hypotheses and twenty-two specific predictions follow from this framework, built explicitly to be tested piece by piece rather than accepted as a single package. Buss and Schmitt also noted, almost in passing, that lifelong monogamy is not how most human societies have actually organized mating: roughly 80 percent of documented societies permit polygyny, serial marriage is common even where it is not, and lifetime sexual exclusivity with one partner is closer to a norm stated than a norm observed [2] — a detail that complicates any reading of the theory as a defense of a particular family arrangement, since the theory was never describing one.

ADVERTISEMENT

The instrument that produced evolutionary psychology’s most cited number

David Buss’s 1989 Behavioral and Brain Sciences target article tested five predictions drawn from this theoretical base against data from 37 samples in 33 countries on six continents and five islands, a total of 10,047 respondents ranging in mean age from 16.96 in New Zealand to 28.71 in West Germany [3]. Two instruments were used: a ranking task asking respondents to order 13 characteristics by desirability in a marriage partner, and a rating task scoring 18 characteristics from 0 (“irrelevant or unimportant”) to 3 (“indispensable”). Each instrument was forward-translated into the local language by bilingual collaborators, back-translated into English by a second bilingual speaker, and reconciled by a third before use, and in most cases “data were collected by native residents within each country and mailed to the United States for statistical analysis” [3] — the study was as much a feat of postal logistics as of psychology, with coded response sheets converging on one American research team from Nigeria, Estonia, Taiwan, and Venezuela inside the same data-collection window. Buss was candid about the resulting limitations: samples could not be treated as nationally representative, rural and less-educated respondents were underrepresented despite some deliberate exceptions such as the South African Zulu and Venezuelan community samples, and sample sizes ranged from 55 in Iran to 1,491 in mainland United States [3].

A single punched data card partway out of a wooden card-catalogue drawer among hundreds of others, a small country-code tab and a faded postal stamp visible on its edge
Figure 1. Coded response sheets from thirty-seven cultures converged on one American data-processing center by international mail — the study's global reach ran, in practice, through a very physical postal and card-punch pipeline.Image prompt and art direction by Brecht Corbeel; generation pending.

The headline results are more textured than their retellings. Women rated “good financial prospect” as more important than men did in 36 of the 37 samples, the sole exception being Spain, where the difference ran in the predicted direction but fell short of significance [3]. Men rated physical attractiveness (“good looks”) higher than women in all 37 samples, significant in 34 of them; the three exceptions — India, Poland, and Sweden — reached significance instead on the parallel ranking measure [3]. The starkest sex difference by far concerned age: men preferred spouses 2.66 years younger than themselves on average, women preferred spouses 3.42 years older, and every one of the 37 samples was significant beyond the .0001 level, the most universal finding Buss reported [3]. He validated the preference data against real demographic yearbooks in 27 countries and found the preferred age gap correlated with the actual gap at marriage at r = .68 for men and r = .71 for women, both p < .001 [3] — a rare instance in this literature where stated preference and a real-world outcome visibly agree, and the largest preferred gaps of all appeared in Nigeria and Zambia, the two samples drawn from societies that practice substantial polygyny.

Two results carried far less cross-cultural agreement, and both are the parts most popular summaries drop. Ambition–industriousness ran in the predicted direction (women valuing it more) in 34 of 37 samples but reached significance in only 29, and reversed direction, significantly, among Zulu respondents in South Africa, where Buss’s own collaborator attributed the reversal to a local division of labor in which “it is considered women’s work to build the house, fetch water, and perform other arduous physical tasks, whereas men often travel from their rural homes to urban centers for work” [3]. Chastity showed even less consistency: only 23 of 37 samples, 62 percent, produced the predicted sex difference at significance, and in Sweden, Norway, Finland, the Netherlands, and West Germany respondents rated a partner’s prior sexual experience as irrelevant or unimportant, with Buss noting that “a few subjects even indicated in writing that chastity was undesirable in a potential mate” [3]. Buss’s own conclusion treated this variability as evidence for his broader thesis rather than against it, arguing that cultural heterogeneity this wide “gives greater credibility to the empirical sex differences that transcend this cultural diversity” [3]. Whether or not that inference holds, the variability itself is real, reported in his own tables, and routinely absent from the two-sentence version of the study that circulates outside the field. So is his own caveat that “male and female preference distributions overlap considerably, in spite of mean differences,” and that in every single sample both sexes ranked “kind-understanding” and “intelligent” above earning power and physical attractiveness [3] — the sex-linked traits that made the study famous were, on Buss’s own numbers, never what either sex wanted most.

A preregistered replication cuts most effect sizes in half without erasing them

Thirty-one years and one replication crisis later, a team led by Kathryn Walter and Daniel Conroy-Beam recruited 14,399 respondents, 54.93 percent female, from 45 countries in 2016, collected the data in person rather than online specifically to preserve representativeness in countries with uneven internet access, preregistered every predictor, moderator, and control variable at the Open Science Framework before running a single analysis, and published both the raw data and the analysis code alongside the paper [4]. This is what auditing a classic finding properly looks like in practice, and the design choices matter as much as the results: multilevel models replacing Buss’s per-country t-tests, a refined five-item rating scale that split Buss’s double-barreled “education and intelligence” item and added a dedicated kindness item, a fuller seven-point response range in place of Buss’s four-point scale, and actual partner age reported alongside ideal preferences as an independent check [4]. To rule out the possibility that cross-cultural variation in the results was itself an artifact of uneven sample sizes, the team plotted each country’s sex difference against its sample size and confirmed the expected funnel shape, in which larger samples cluster near the average effect and smaller ones scatter more widely — a sign that some of the apparent variability across countries in this literature is ordinary sampling noise rather than a real cultural signal [4].

A rugged field tablet in a padded case showing a seven-point rating scale mid-tap, resting on a folding table beside a paper preregistration printout with a visible timestamp
Figure 2. The 2016 replication traded mailed paper forms for tablets carried into forty-five countries in person — a design choice built specifically to avoid the online-panel bias that skews toward wealthier, more connected respondents.Image prompt and art direction by Brecht Corbeel; generation pending.

The central sex differences held. Women rated financial prospects higher than men on a standardized scale comparable to Cohen’s d, b = −0.30, p < .001; men rated physical attractiveness higher, b = 0.27, p < .001; and the age gap replicated almost exactly, with men’s actual partners averaging 2.26 years younger and women’s partners averaging 2.43 years older, both p < .001 [4]. Smaller, more consistent sex differences also emerged for kindness, intelligence, and health, each favoring women’s stated preference by roughly a tenth of a standard deviation, echoing Buss’s original finding that both sexes valued these traits highly while still showing a slight sex-typed edge in exactly how highly [4].

ADVERTISEMENT

What moved was magnitude. Combining all five preferences into a single multivariate distance between the sexes, a technique unavailable to Buss’s per-trait t-tests, produced an overall Mahalanobis D of 0.73 across the 45 countries, ranging from 0.30 in Nigeria to 1.42 in Georgia, and the team stated the historical comparison directly: “Using the data from all 18 preferences (excluding age) from Buss (1989), Conroy-Beam et al. found the mean Mahalanobis D between males and females to be 1.46” [4]. On directly comparable multivariate terms, the overall sex difference in mate preferences was roughly half the size it had appeared to be thirty years earlier. A separate univariate reanalysis of Buss’s own 1989 data using 19 preference items had gone further still, reporting D = 2.41 and classifying respondent sex from preferences alone with 92.2 percent accuracy [9] — a number that made the sexes sound almost non-overlapping. Walter and colleagues’ own cross-validated classifier, trained and tested on the new sample, achieved 63 percent average accuracy [4]: above chance, unmistakably real, and a long way from 92 percent. The multivariate framing that once made mate preferences look like one of the largest psychological sex differences on record is the same framing that, applied honestly to newer, preregistered data, now supports a considerably more modest claim.

One hypothesis failed outright rather than merely shrinking. Pathogen prevalence, the ecological variable Steven Gangestad and Buss had proposed in 1993 as a driver of cross-cultural variation in preferences for attractiveness, health, and intelligence, predicted only a preference for financial prospects in the new data, and even that relationship disappeared once latitude, GDP, world region, and religion were entered as controls; the team reported plainly that their results “did not replicate the findings of Gangestad and Buss (1993) or Gangestad et al. (2006)” [4].

Gender equality moves one preference and not the others, and that split is the finding

The rival account of these same numbers is social role theory, proposed by Alice Eagly and Wendy Wood in 1999, which locates the origin of sex-typed mate preferences not in evolved psychology but in the contemporaneous division of labor between providers and homemakers [5]. Eagly and Wood’s specific, falsifiable prediction was that sex differences tied to that division of labor should shrink as gender equality rises, and their reanalysis of Buss’s 37-culture dataset, using the United Nations’ 1995 Gender Empowerment Measure and Gender-Related Development Index, found real support for it, but only for some preferences. The sex difference in valuing a spouse’s earning capacity correlated negatively with the Gender Empowerment Measure, r = −.43 on the ranking version of the item, p < .05, and the sex difference in valuing a spouse’s skill as “a good housekeeper and cook” correlated even more strongly, r = −.62, p < .001: as gender equality rose, men lost interest in domestic skill faster than women gained interest in earning power, but both moved toward each other [5]. The sex difference in preferred spousal age also shrank as gender equality rose, with the text reporting explicitly that “as gender equality increased, women expressed less preference for older men, men expressed less preference for younger women, and consequently the sex difference in the preferred age of mates became smaller” [5]. But no comparably strong or consistent relationship emerged for physical attractiveness in Eagly and Wood’s own data — their theory, on their own numbers, explained the resource and domestic-skill preferences convincingly and the attractiveness preference weakly at best.

Marcel Zentner and Klaudia Mitura extended the test in 2012 with a purpose-built ten-country sample of 3,177 respondents and a thirty-one-nation follow-up of 8,953 volunteers, using an updated gender-equality measure and a composite sex-difference score, and reported the same direction of effect across both samples: preference sex differences ran largest in the least gender-equal nations and smallest in the most equal ones [6]. Zentner’s own framing of the finding resisted the simplest anti-evolutionary reading of it: he suggested that the capacity to shift mating psychology quickly in response to social change “may itself be driven by an evolutionary program that rewards flexibility over rigidity” [6] — a claim that a rapid, socially responsive preference shift and an evolved mechanism are not necessarily opposites.

An open ring binder showing two facing columns of translated survey wording, a red proofreading mark circling a mismatched phrase between the forward and back translation
Figure 3. Every one of the study's forty-five language versions had to be forward-translated, back-translated, and reconciled before a single response could be compared across countries — the quiet infrastructure beneath every cross-cultural number in this literature.Image prompt and art direction by Brecht Corbeel; generation pending.

Then a larger, preregistered test came in against the broader version of the moderation claim. Walter and colleagues tested five separate gender-equality indices, including the Gender Empowerment Measure and Global Gender Gap Index used by the two prior studies, against sex differences in all five mate preferences, and reported that “gender equality did not robustly predict sex differences in any of the mate-preference measures,” with a single partial exception: the Global Gender Gap Index predicted the sex difference in financial-prospects preference, b = 0.06, p = .036, replicating one piece of Zentner and Mitura’s finding and nothing more [4]. What gender equality did predict, robustly and in both this study and Eagly and Wood’s much earlier one, was the shrinking sex difference in actual partner age, the one variable in the whole battery that describes a real-world outcome rather than a stated ideal [4]. A fully independent test closed off the most likely objection to this whole pattern: Lingshan Zhang and colleagues surveyed 3,073 respondents in 36 countries using Buss’s original 1989 items verbatim, specifically to rule out the possibility that Walter’s redesigned five-item scale had itself caused the non-replication, and found the identical picture — real, significant sex differences in attractiveness and earning-capacity preference, and, in their own words, “little support for the social roles account of sex differences in mate preferences” [11]. The moderation hypothesis did not fail uniformly across the board. It failed specifically for the two preferences, attractiveness and financial prospects, that carry the most cultural weight in this debate, and held specifically for the one variable, age, that touches lived behavior rather than a stated ideal.

Two standing theories, and the moderation evidence that cuts against both simplifications

Neither side of this dispute gets to declare the domain-specific pattern a clean win. Pure social role theory predicts that a sex difference tracking a social division of labor should erode wherever that division erodes, and by that standard the theory has a real problem with attractiveness: postindustrial gender equality has advanced further and faster on nearly every measured axis than it had in 1999, and the male preference for physical attractiveness has not correspondingly shrunk in any of the three independent tests that have since looked for it [5, 4, 11]. Pure evolutionary psychology has an equally real problem with age: if the preference for a fertility-linked partner-age gap is a species-typical adaptation calibrated to ancestral reproductive value, it should not be tracking a country’s contemporary gender-equality score at all, and it does, in the same direction, in two independently collected datasets three decades apart [5, 4].

ADVERTISEMENT

What the multivariate reframing from Daniel Conroy-Beam’s research group adds is a reason the two underlying questions, how big is any one sex gap and how do several gaps combine, can have genuinely different answers. People do not evaluate potential partners one trait at a time; they weigh a whole pattern of traits together, so that a modest deficit on one dimension can be offset by a surplus on another, the way two towns 100 miles apart on each of two separate axes end up not 100 but roughly 141 miles apart along the diagonal between them [9]. In simplified form, treating the dimensions as uncorrelated and standardized, that compounding looks like an ordinary Euclidean distance across k traits:

Dd12+d22++dk2 D \approx \sqrt{d_1^2 + d_2^2 + \cdots + d_k^2}

The Mahalanobis distance Conroy-Beam’s team actually computed additionally weights each dimension by the sample’s covariance structure, but the intuition survives the simplification: several individually modest sex differences on correlated traits compound into a larger difference in the overall pattern, which is exactly why the same body of evidence can support both “the univariate sex gaps are considerably smaller than people assumed” and “the pattern-wise sex difference remains substantial” without contradicting itself [9]. A separate companion analysis of the 45-country sample tested eight competing computational models of how people actually combine several preferences into one evaluation of a potential partner — including a compensatory Euclidean-distance model, a simple additive weighting model, an aspiration-threshold model that treats each trait as a minimum bar to clear, and a curvilinear model, alongside two null controls — and found the Euclidean, compensatory model fit best across every one of the 45 countries tested [10]. That result, that people trade traits off against each other rather than screening candidates item by item, is the closest thing this literature currently has to an uncontested cross-cultural universal, and it is a different kind of claim from any statement about the size of a particular sex gap.

A steel archive cabinet with two labelled drawers, an older one and a newer one, a key caught mid-turn in the newer drawer's lock with the drawer just cracked open
Figure 4. Two datasets collected three decades apart now sit in the same archive, and the argument between evolved-disposition and social-role accounts of mate preference is, in the end, an argument about how to read the distance between those two drawers.Image prompt and art direction by Brecht Corbeel; generation pending.

What people ask for is not what moves them in the room

A second complication sits underneath all of the cross-cultural numbers and has nothing to do with culture at all: stated preferences and revealed choices are measured differently, and in the most careful tests they do not agree. Paul Eastwick and Eli Finkel ran a full speed-dating study at Northwestern University in which participants first stated their ideal-partner preferences on a questionnaire, then attended a two-hour event with nine to thirteen four-minute dates each, generating 206 matched pairs, then reported their actual romantic interest in each partner and in any other person they became interested in over the following month, adding 143 further “write-in” targets to the dataset [7]. The stated preferences looked exactly like Buss’s: a medium-to-large sex difference favoring men’s stated emphasis on physical attractiveness, mean d = −.55, and a small-to-medium difference favoring women’s stated emphasis on earning prospects, mean d = .35, with no sex difference at all on a comparison trait, personableness, that neither theory expects to be sex-typed [7]. But when the same participants’ actual romantic interest in their matches and write-ins was regressed on those partners’ attractiveness and earning prospects, the sex difference vanished. The correlation between a partner’s attractiveness and romantic interest in them was .43 for men and .46 for women, a nonsignificant difference, r = .03, p = .673; the correlation for earning prospects was .19 for men and .16 for women, again nonsignificant, r = −.04, p = .480 [7]. A handful of individual outcome measures did show scattered sex differences — men, for instance, were somewhat more likely than women to initiate contact with a match they had rated attractive — but across 25 paired regression comparisons for each trait, only two to three showed a significant sex difference in either direction, well within what chance alone would produce [7]. On the traits that made Buss’s study famous, people’s stated ideals and the traits that actually predicted whom they wanted turned out to be two different constructs.

Two paper rating forms side by side on a desk, an ideal-partner questionnaire and a live-interaction record sheet, a fountain pen resting with its nib still touching a freshly marked box on the second sheet that contradicts the mark on the first
Figure 5. The same person's stated ideal and their in-the-moment rating of an actual encounter do not have to agree, and in the largest tests of the question, they mostly do not.Image prompt and art direction by Brecht Corbeel; generation pending.

Eastwick returned to the question in 2014 with Laura Luchies, Finkel, and Lucy Hunt, meta-analyzing 97 studies spanning three stages of relationship formation, from hypothetical targets through live first encounters to established relationships, and found that the single-study result generalized across the wider literature. Physical attractiveness predicted romantic evaluation with a moderate-to-strong effect, r ≈ .40, for both sexes; earning prospects predicted romantic evaluation with a small effect, r ≈ .10, for both sexes; and the sex difference in either correlation was small and, across the full set of studies, “uniformly nonsignificant,” with a difference between the male and female correlations of only r = .03 [8]. The 2008 team framed the gap between what people say they want and what actually predicts their romantic response through Richard Nisbett and Timothy Wilson’s older argument that people frequently lack accurate introspective access to the real causes of their own judgments [7]. If that framing is right, it does not mean the cross-cultural literature above has been measuring nothing. It means that literature has been measuring a real, stable, and now well-replicated thing — just not quite the thing its questionnaire item names.

This is what asking psychology to grow up after 2011 actually produced

Every methodological complaint the field’s post-2011 replication reckoning raised against this literature has a named fix in the studies above, and each fix was actually applied rather than merely promised. Small, unrepresentative, online-convenience samples: Walter and colleagues collected 14,399 responses in person across 45 countries specifically to avoid the internet-access bias that skews online panels toward wealthier, more educated respondents in lower-income countries [4]. Undisclosed analytic flexibility, the specific failure mode named in the 2011 paper that helped trigger this reckoning across psychology: the predictors, moderators, and controls in the 45-country study were preregistered at the Open Science Framework before the data were analyzed, and the raw dataset and code were published alongside the paper rather than made available “on request” [4]. A single research group’s idiosyncratic instrument quietly driving a result: Zhang and colleagues deliberately re-ran the gender-equality test using Buss’s original 1989 wording, item for item, to separate a measurement artifact from a genuine non-replication, and found the same null outcome either way [11]. Competing theoretical camps each running studies designed to confirm their own priors: Walter’s team, drawing on more than seventy co-authors based in psychology departments across every inhabited continent, built one preregistered analytic framework that tested the evolutionary and social-role predictions against the identical data at the same time, rather than each camp separately reanalyzing whichever dataset suited it best [4].

A dated preregistration printout with a fresh red date-stamp lying across a retired stack of 1980s paper coding sheets
Figure 6. A timestamped preregistration filed before the data existed is the single procedural change that separates this literature's newest results from its oldest ones.Image prompt and art direction by Brecht Corbeel; generation pending.

The result is a literature that resembles neither camp’s caricature of it. It is not true that mate preferences are culturally arbitrary and only appear sex-differentiated because of patriarchal role assignment: the age preference, and on the best current multivariate evidence the overall pattern of preferences, are real, replicated across independent samples spanning three decades, and not fully explained away by any gender-equality score yet measured [4, 9]. It is equally not true that mate preferences are fixed, universal, sex-linked adaptations immune to social conditions: the age-gap sex difference tracks gender equality in two independently collected datasets, the pathogen-prevalence account failed a preregistered test outright, and the multivariate sex difference that once looked nearly total, 92.2 percent classification accuracy, now looks real but partial, 63 percent [5, 4]. What survived thirty-one years and a considerably stricter methodology was smaller than either side’s opening claim, and a good deal more interesting than a tie.

Distributions are not norms, and the difference has to be stated on purpose

A careful reader can take a specific, bounded set of conclusions from this record. A small number of sex-typed differences in stated mate preference, chiefly age, physical attractiveness, and financial prospects, are real, have now been observed in independently collected samples spanning more than fifty countries and three decades, and are not an artifact of any single research group’s instrument. Their size, measured the way that matters for describing an actual population rather than headlining a summary, is modest: multivariate distance estimates have moved from claims of near-total separation between the sexes’ preference profiles toward substantial overlap, depending on which and how many traits are pooled together [9, 4], and Buss’s own 1989 data showed considerably overlapping distributions from the very start [3]. Whether the underlying mechanism is an evolved, calibrated disposition or an ongoing social accommodation to role expectations remains a live and only partially resolved question, and the honest answer differs depending on which preference is being asked about: the evidence increasingly favors an evolved-and-largely-stable account for attractiveness and financial prospects, and a more socially responsive account for partner age, the one variable where every available test has found gender-equality moderation that the other preferences did not show [5, 4]. And underneath the entire cross-cultural question sits a separate, unresolved measurement problem: what a person says they want on a questionnaire and what actually predicts whom they pursue in a live interaction are, across a 97-study meta-analysis, measurably different quantities, with no reliable sex difference between them at all [8].

None of this literature, read accurately, licenses a claim about what any particular person should want, whom any particular person should choose, or how a relationship or a society ought to be arranged. Every number reported above describes the mean and the spread of a self-reported distribution across a sample of people who were asked to report a preference, not to justify one; a population-level average with heavy overlap between groups says very little that is reliable about any individual inside either group, and a well-replicated “is” about a stated-preference distribution has never, by itself, supplied an argument for an “ought” about how anyone must live or whom anyone must choose. The study of mate preference has spent thirty-seven years, tens of thousands of respondents, and one serious methodological reckoning arriving at a smaller, more qualified, and considerably more interesting answer than the one that made the original finding famous. That is what the evidence, read carefully, supports. It is also the full extent of what it supports.