A story that costs nothing to tell is not thereby true

An evolutionary claim about the mind is one of the cheapest hypotheses a person can produce. State a behavior, state a plausible ancestral problem it might have solved, assert that the behavior therefore evolved to solve it, and the story is finished. It requires no data collection, no comparison group, no falsifiable prediction stated in advance. It can be assembled in the time it takes to order a coffee, and it is assembled that way constantly, in newspaper science sections and dinner-table arguments alike, wherever someone wants to explain why people do the thing they just noticed people doing.

This is not a rhetorical exaggeration invented for this article. It is close to the literal charge Stephen Jay Gould and Richard Lewontin brought against evolutionary biology’s mainstream in 1979, in a paper that gave the field a vocabulary it has never fully shaken off. Their opening line is direct about its target: “An adaptationist programme has dominated evolutionary thought in England and the United States during the past 40 years. It is based on faith in the power of natural selection as an optimizing agent” [1]. The programme they describe proceeds by breaking an organism into separate traits and inventing an adaptive story for each one, and its central intellectual sin, in their account, is treating plausibility as if it were evidence.

Gould and Lewontin needed an image sturdy enough to carry that argument, and they found it in architecture rather than biology. The dome of St Mark’s Basilica in Venice rests on four rounded arches, and where two arches meet at a right angle they leave behind a tapering triangular gap. “Spandrels — the tapering triangular spaces formed by the intersection of two rounded arches at right angles — are necessary architectural by-products of mounting a dome on rounded arches. Each spandrel contains a design admirably fitted into its tapering space,” they wrote, going on to describe the evangelist mosaic filling each one [1]. Here is their point compressed into one artifact: the space exists purely as a structural consequence of the arches, not because anyone wanted a triangular gap; the elaborate mosaic that fills it arrived afterward, fitted to a shape that was never designed with it in mind. Mistake the byproduct for the intended feature, and you would write a very confident essay about why cathedral architects evolved a preference for tapering triangular evangelist niches. The essay would be wrong about causation while being locally, superficially correct about fit.

ADVERTISEMENT

Their sharpest illustration of the same error in biology has nothing to do with spandrels and everything to do with a joke that still lands almost fifty years later: “We fault the adaptationist programme for its failure to distinguish current utility from reasons for origin (male tyrannosaurs may have used their diminutive front legs to titillate female partners, but this will not explain why they got so small)” [1]. A trait can be used for something today without that use being the reason it exists. The habit of asking only “what is this good for” and stopping there, satisfied, is the whole of what they are attacking — not adaptation as a real evolutionary process, which neither author denied, but adaptation as a reflex explanation applied to every trait without checking whether a non-adaptive account fits the facts better. Their own list of alternatives is worth stating in full, because it is longer and more concrete than the caricature the paper is sometimes reduced to: random fixation of alleles, non-adaptive byproducts of developmental correlation with a selected trait (they name allometry, pleiotropy, and mechanically forced correlation specifically), the separability of adaptation from the selective process that produced it, multiple adaptive peaks that selection could have settled on for reasons having nothing to do with optimality, and current utility as nothing more than an accident riding on a structure built for another reason entirely [1]. Their proposed corrective was not to abandon adaptation but to return to Darwin’s own pluralism, and to treat the null hypothesis — this trait may not be for anything in particular — as one live possibility among several rather than the option adaptationists could skip past.

The paper did not go unanswered, and the reply worth taking seriously is not the one that simply restated adaptationism louder. David Queller’s 1995 response, tellingly titled in parody of the original, called Spandrels “an opinion piece, a polemic, a manifesto, and a rhetorical masterpiece” even while attacking it [3] — a concession that the criticism landed as rhetoric before it landed as science, and that its persuasive force sometimes outran its evidentiary content. The substantive version of the adaptationist reply, developed across the following decades by researchers working the problem empirically rather than polemically, is that Gould and Lewontin were right to demand a stopping condition on plausible storytelling and wrong to suggest that the field, done properly, lacks one.

Confer and colleagues’ 2010 defense in American Psychologist is the field’s own fullest answer to the charge, and it argues its case with refutations rather than assurances. They point to the kin-altruism theory of male homosexuality, which predicted that gay men should direct measurably more generosity toward genetic relatives than straight men, as a hypothesis the field actually gave up on: a direct test matching homosexual and heterosexual men for age, education, and ethnicity found no support for any of its key predictions, and Confer and colleagues state plainly that “the kin altruism theory of male homosexuality has been refuted” [11]. They set that failure beside cases that survived equally specific tests — a memory-encoding hypothesis that would have been falsified had survival-relevant material failed to show a recall advantage, and did not; an “auditory looming bias” and a “commitment skepticism bias” that both generated predictions capable of coming out negative and did not [11]. Their summary judgment is the line worth holding onto through the rest of this article: “The sometimes reflexive charge that evolutionary psychological hypotheses as a rule are mere ‘just-so stories,’ however, is simply erroneous, as the examples above demonstrate” [11]. That is a defensible claim about the field’s better work, not a blanket acquittal — which is exactly why the toolkit below, rather than either side’s rhetoric, has to do the remaining work of telling the two kinds of case apart. The rest of this article is an inventory of the actual stopping conditions available — the tools that turn “here is a story that fits” into “here is a story that predicted something specific and turned out to be right, or didn’t.”

An operant conditioning chamber pulled from its cabinet with its side access panel removed, a wired response lever's connector caught only half-seated on its terminal beside a second, empty mounting bracket cast into the same frame but never fitted with a lever
Figure 1. Not every part of a working instrument is there because it does something: the bracket cast into this chamber's frame never held a lever, and it is as good a model of a spandrel as anything in San Marco.Image prompt and art direction by Brecht Corbeel; generation pending.

Four questions, one grid, and the confusion that comes from collapsing them

The single most useful piece of intellectual furniture in this whole field predates evolutionary psychology as a named discipline by three decades. Niko Tinbergen, writing in 1963 to honor Konrad Lorenz, set out to clarify what a complete biological explanation of a behavior actually requires. Julian Huxley, he noted, had already proposed three: “the three major problems of Biology: that of causation, that of survival value, and that of evolution — to which I should like to add a fourth, that of ontogeny” [2]. Those four became, in the language the field now uses, mechanism (the proximate machinery that produces the behavior right now), function (the fitness consequence that made the behavior favored by selection), phylogeny (the evolutionary history and ancestry of the trait), and ontogeny (how the trait develops across an individual’s lifetime).

The grid matters because these are four separate questions with four separate kinds of evidence, and a correct answer to one is not evidence for a particular answer to another. This is exactly the confusion sitting underneath the tyrannosaur joke: an answer about current utility (function, loosely construed) was being mistaken for an answer about phylogenetic origin. It recurs constantly in lay discussion of evolutionary psychology, usually in a more mundane form than dinosaur forelimbs. Take a claim as ordinary as the observation that men are, on average, more physically aggressive than women. A mechanistic answer describes the proximate machinery — circulating androgens, threat-appraisal circuitry, situational triggers. A functional answer describes what aggression accomplished for fitness in some set of ancestral conditions — resource or mating competition, status defense. An ontogenetic answer describes how the sex difference emerges across development — prenatal hormone exposure, childhood socialization, or some interaction of both. A phylogenetic answer places the pattern in a comparative context — where it sits relative to the male intrasexual competition seen broadly across mammals with polygynous mating systems. All four can be true simultaneously and none of them stands in for the others. The mechanism explains how, not why it exists; the function explains why it was favored, not how it is produced moment to moment; and neither one, by itself, tells you the phylogenetic depth of the trait or the developmental pathway that builds it in a given individual. A popular article that reports a hormone correlation and calls it “the evolutionary explanation” has usually confused the first question for the third, and a reader equipped with Tinbergen’s grid can see the seam where that happened.

ADVERTISEMENT
A habituation-stage infant eye-tracker mount in a bright testing alcove, its camera boom caught mid-lower and a blank fixation-target card only halfway seated in its holder with one corner still lifted
Figure 2. The rig measures where an eye lingers, not why it lingers there: mistaking that reading for an answer about function rather than mechanism is the single most common way Tinbergen's four questions get collapsed into one.Image prompt and art direction by Brecht Corbeel; generation pending.

Special design is the strongest test the toolkit has, and the easiest one to counterfeit

If Tinbergen’s grid tells you which question is being asked, George Williams’s 1966 book, Adaptation and Natural Selection [4], supplied the sharpest available standard for answering the functional question well. As David Schmitt and June Pilcher summarize it, “Williams (1966) provided perhaps the most influential and enduring guide to identifying historical adaptations. He argued that only when an attribute shows evidence of special design for the purpose of increasing fitness should one consider an attribute to be an adaptation” [5]. Two features have to be present together, on this standard: the trait has to plausibly increase fitness, and it has to show special design — meaning it looks like a built solution to a specific problem rather than a generic capacity or an accident. Schmitt and Pilcher’s own gloss on what special design looks like is concrete: “if an attribute is extremely efficient, subtly complex, or incredibly specialized, or emerges universally in all members of a species, then one can think of the attribute as possessing design specificity,” and further, “adaptations are expected to be efficient or economical, in the sense that little that is energetically wasteful is retained in an adaptation’s structure over evolutionary time” [5]. Economy, efficiency, precision, reliability, and functional specificity are the fingerprints Williams asked researchers to look for — not because a designer left them, but because selection, acting over enough generations against enough failed alternatives, tends to leave a structure that looks purpose-built even though nothing purposeful built it.

The value of the standard is that it rules things out. A capacity that is general-purpose, imprecise, wasteful of energy, or indifferent to the specific structure of the problem it is supposedly solving fails the test regardless of how good the accompanying story sounds. One of the cleanest demonstrations of a trait that passes comes from outside evolutionary psychology proper, in classic conditioning research: rats and other animals exposed to gamma radiation, and made ill hours afterward, reliably develop a strong aversion specifically to the flavor of food or water they had consumed before falling ill — not to the light, the sound, the cage, or any of the other incidental cues present at the same time [13]. General associative-learning theory, in its plainest form, predicts that any sufficiently salient cue should be equally associable with any consequence; it has no principled reason to expect taste specifically, rather than light or sound, to bond so strongly to illness across a delay of hours that would defeat ordinary conditioning entirely. A toxin-avoidance mechanism built specifically to link flavor to gut consequence, and built to tolerate exactly the kind of delay that food poisoning actually involves, predicts precisely this pattern and nothing else. That selectivity — this cue bonds to that consequence, and no other combination works nearly as well — is what special design means in practice, and it is why the finding has stood as a foundational case for “prepared,” domain-specific learning ever since.

The trap sits immediately on the other side of the same vocabulary. Special-design language is easy to borrow without doing special-design work: describe any trait as “efficient” or “precisely tuned” after the fact, without a prediction that could have come out otherwise, and the terminology dresses up exactly the kind of unfalsifiable storytelling Gould and Lewontin attacked, now wearing Williams’s own words as camouflage. The discipline that separates a genuine special-design argument from a counterfeit one is whether the specific, structured prediction was stated before the confirming data were in hand, and whether a design that failed to show that structure — a generalized fear response instead of a snake-specific one, a taste-blind conditioning bias instead of a taste-specific one — was a live possibility the researcher would have had to accept. A special-design claim that cannot specify in advance what pattern would have falsified it has not used Williams’s criterion; it has quoted it.

A paired conditioned-taste-aversion testing rig on a bright bench, a single drop hanging unfallen from the near chamber's drip-bottle spout while the far chamber's tubing coupling sits open and disconnected
Figure 3. Special design is the strongest evidence the toolkit has, because a mechanism this selective — taste bound tightly to illness, almost nothing bound to shock — is hard to explain except as a built solution to a specific problem.Image prompt and art direction by Brecht Corbeel; generation pending.

Universal is not the same claim as evolved

The next tool in the kit asks a simpler-sounding question: does the trait show up in every human population researchers have checked, or only in some? Universality across cultures is one of the eight broad categories of evidence Schmitt and Pilcher catalogue in their framework for evaluating adaptation claims, sitting alongside psychological, physiological, medical, genetic, phylogenetic, and hunter-gatherer evidence as one leg among several [5]. The intuition behind it is reasonable: if a behavior or preference appears in societies with no contact and no shared cultural transmission, cultural invention becomes a strained explanation, and a species-typical, genetically canalized developmental program becomes more plausible.

The trouble is that universality is compatible with more than one underlying cause, which is exactly why it can only ever be one leg of a larger argument rather than a verdict on its own. A trait can be universal because it is a genetically specified adaptation. It can also be universal because every human environment reliably presents the same problem, and every human developmental system reliably re-derives the same solution through learning rather than inheriting it outright — a fear of unlit spaces, say, could reflect an evolved threat-detection bias, or it could reflect the mundane fact that every human child eventually learns, in whatever culture raises them, that visibility correlates with safety. And a trait can be universal for the same reason a spandrel is universal in every domed building built the same way: because it rides along, unavoidably, with something else that actually was selected. Universality tells you the trait is not a local cultural quirk. It does not, by itself, tell you which of these three stories is correct, and treating it as if it settles the question is the underdetermination this tool is prone to.

The point cuts in an unexpected direction with a well-known test from comparative and developmental psychology. Mirror self-recognition testing — marking a subject and watching whether they use a mirror to investigate the mark on their own body — is often treated as a benchmark of a species-typical developmental milestone in young children. Cross-cultural work on exactly this test found sharply uneven results: one 2010 study found only about three percent of Kenyan children and none of the Fijian children tested passed a mark-test measure that the large majority of children in Western samples pass by a comparable age, a gap best explained by differences in cultural practice around mirrors and self-inspection rather than by any difference in underlying cognitive development [12]. A measure researchers had reason to expect would generalize turned out to vary enormously by culture instead — which is the universality tool’s failure mode demonstrated from the inside, on a trait many people would have guessed was about as universal as they come.

ADVERTISEMENT
An open flat-file drawer of stimulus card sets for cross-cultural forced-choice testing, blank card backs standing on edge in dividers, one card pulled half out of its kraft-paper sleeve
Figure 4. A pattern that turns up in every culture surveyed is evidence worth having, but it underdetermines the question: a universally reinvented lesson looks the same in this drawer as a universally inherited one.Image prompt and art direction by Brecht Corbeel; generation pending.

Shared machinery across relatives is harder evidence than a pattern in one species

A related but distinct tool looks past culture entirely and asks whether the trait, or something recognizably like it, appears across species related by descent, in a pattern that tracks the branching structure of the evolutionary tree rather than the map of human trade routes. Schmitt and Pilcher list this as its own category — phylogenetic evidence, drawing on animal ethology, comparative psychology, primatology, physical anthropology, and paleontology [5] — because a trait shared with close relatives and absent or different in more distant ones is much harder to explain as coincidence or cultural transmission than a trait found only within one species. If a capacity is inherited from a common ancestor (homologous) rather than independently reinvented (convergent) or culturally learned, the comparative pattern across the phylogenetic tree should show it, in the same way shared vocabulary and shared grammar show inheritance between languages descended from a common source.

Mirror self-recognition is again the clearest worked case, this time run as a comparative test across species rather than across cultures. Gordon Gallup’s original 1970 design anesthetized chimpanzees, marked them with odorless dye in a spot only visible in a mirror, and found that mark-directed touching rose sharply once a mirror was available [12]. The pattern that followed decades of replication attempts across species is genuinely comparative: chimpanzees, bonobos, and orangutans pass reliably; most monkey species fail reliably; and a scattering of more distantly related cases — some elephants, some cetaceans, and a contested finding in magpies that a 2020 replication failed to reproduce — sit awkwardly outside any clean story about great-ape-specific cognition [12]. That pattern, uneven as it is, is the kind of evidence this tool is built to supply: a trait’s distribution across a phylogeny that at least roughly, if imperfectly, tracks relatedness to humans.

It is also, on its own, not proof of what the trait means. Daniel Povinelli’s long-standing critique of the paradigm argues that passing the mark test may show only that an animal can reason about the causal contingency between its own movements and the mirror image — a general capacity for detecting “an odd, controllable entity” — without that capacity amounting to self-awareness in the richer sense the test is usually taken to demonstrate [12]. A strong comparative pattern narrows the space of explanations. It does not, by itself, tell you which specific psychological trait the pattern is evidence for, which is why the toolkit’s tools are meant to be used together rather than any one of them treated as dispositive.

Heritability tests the wrong thing, and a fixed adaptation should score low on it

Here is the tool in the kit that runs backward from intuition, and it is worth a section of its own precisely because the intuitive reasoning is so widely shared and so wrong. The common assumption is that if a psychological trait is heritable — if twin and family studies find that genetic differences between people predict differences in the trait — that counts as evidence the trait is an evolved adaptation. The counterintuitive result, laid out by John Tooby and Leda Cosmides in 1990, is closer to the opposite: a trait that is a genuinely universal, tightly engineered adaptation should show low heritability, precisely because it is universal and tightly engineered.

Their argument runs through a specific mechanism. Complex adaptations require many genes working together in a coordinated way to build functioning machinery, and sexual recombination shuffles gene combinations every generation, which makes it statistically improbable that a well-functioning version of a complex, many-gene system would persist as a segregating genetic difference between individuals rather than becoming the shared, standard-issue design nearly everyone inherits. As they put it: “Selection, interacting with sexual recombination, tends to impose relative uniformity at the functional level in complex adaptive designs, suggesting that most heritable psychological differences are not themselves likely to be complex psychological adaptations. Instead, they are mostly evolutionary by-products, such as concomitants of parasite-driven selection for biochemical individuality” [7]. Applied to personality research specifically, their conclusion is blunt: “most heritable personality differences are not the expression of different adaptive strategies. They are either mutationally driven genetic noise, or else an incidental by-product of an adaptation that has nothing to do with personality per se—pathogen-driven selection for biochemical diversity” [7].

The logic generalizes cleanly once it is stated once. Nobody disputes that having two eyes, or the basic capacity to acquire a native language, is a genuine adaptation — and neither trait shows meaningful heritable variation in the twin-study sense, because selection has driven the underlying machinery so close to fixation that almost everyone gets a working copy. What does show up as heritable variation is disproportionately the material selection has not (yet, or ever) cleaned up: genetic noise, frequency-dependent variation maintained by shifting selective pressures, or byproducts of selection acting on something else entirely, such as pathogen resistance. A finding that some psychological trait is substantially heritable is therefore not, by itself, evidence that the trait is an ancient, core adaptation. It can just as easily be evidence of the reverse — that the trait sits closer to the noise selection has left behind than to the machinery selection has finished building. This is exactly why the “genetic evidence” leg of Schmitt and Pilcher’s framework has to be handled carefully: heritability is real, measurable, and easy to report as though a high number settles the adaptation question, when in the specific case of traits near fixation the informative signal runs the other way.

A hypothesis that survived: pregnancy sickness as toxin avoidance, checked at all four levels

The clearest way to see the toolkit working is to run an actual hypothesis through it, and pregnancy sickness is among the field’s better-supported cases. Nausea and vomiting in pregnancy afflicts a large majority of pregnant women, concentrated in the first trimester, and the evolutionary hypothesis associated with Margie Profet and developed most thoroughly by Samuel Flaxman and Paul Sherman treats it not as an unfortunate hormonal side effect but as a facultative defense: a mechanism that lowers exposure to plant toxins and foodborne pathogens specifically during the developmental window in which a fetus’s organs are being formed and are least able to tolerate chemical disruption.

Flaxman and Sherman’s 2000 synthesis in the Quarterly Review of Biology pulled together fifty-six studies covering roughly seventy-nine thousand pregnancies across sixteen countries [6], which is the kind of aggregated, multi-population evidence base a plausible story rarely bothers to assemble and a tested hypothesis needs. Run through Tinbergen’s four questions, the case holds together at every level rather than resting on one convenient fact. On ontogeny, the timing is itself part of the evidence: incidence rises through weeks nine to fourteen, when sixty to seventy percent of women experience nausea, and symptoms peak precisely between week six and week eighteen of pregnancy — “when embryonic organogenesis (organ development) is most susceptible to chemical disruption” [6]. The mechanism switches on and off tracking a developmental window rather than running constantly across the whole pregnancy, which is what a targeted defense, rather than a generic hormonal malaise, would predict. On mechanism, the proximate route runs through a sharply heightened sensitivity to smell and taste in the first trimester, widely associated with the hormonal shifts that also define that window, giving the aversion system a plausible physiological trigger tied to the same period.

On function, the design specificity Williams’s standard asks for is present in the pattern of what gets avoided, not just in the existence of nausea generally: “the most-observed aversion was to meats, fish, poultry and eggs,” historically the foods most likely to carry parasites and pathogens dangerous to a compromised or developing system, alongside “strong-tasting vegetables, as well as alcoholic and caffeinated beverages” [6] — several of which carry plant toxins or teratogenic compounds rather than infectious risk. A generic nausea response indifferent to food content would not produce this pattern; a system tuned to the specific hazard profile of early pregnancy would. And the fitness-relevant outcome, the piece Williams’s criterion also demands, is in the data directly: “women who experience morning sickness are significantly less likely to miscarry than women who do not. Women who vomit are significantly less likely to miscarry than those who experience nausea alone” [6]. On phylogeny, the case connects to a broader, independently documented pattern rather than standing alone: it draws on the same class of prepared, taste-linked defense against ingested toxins demonstrated in classic conditioning research across other species [13], suggesting pregnancy sickness recruits and intensifies a toxin-avoidance system with a much older evolutionary pedigree rather than inventing one from nothing.

None of this makes the case closed beyond dispute. Researchers still debate how much of the correlation between morning sickness and lower miscarriage reflects the protective behavior itself, as opposed to a shared upstream cause — placental hormone output robust enough to both signal a viable pregnancy and drive stronger nausea — running underneath both. That is a real, open question about mechanism, not a reason to discard the hypothesis; it is the normal condition of a functioning research program, which is exactly what distinguishes this case from the ones the rest of this article treats as settled either way.

An empty two-bottle preference-test cage on a bench, a tinted bottle's cap caught only part-threaded and crooked while the matched clear bottle beside it sits fully capped and ready
Figure 5. Pregnancy sickness passes the test this rig was built to run: the aversions it produces are not random, they are aimed at exactly the foods most likely to carry the toxins and pathogens a first-trimester embryo cannot yet afford.Image prompt and art direction by Brecht Corbeel; generation pending.

A hypothesis that didn’t survive contact with preregistration: the ovulatory-shift literature

The toolkit’s value is clearest when it is also shown failing a hypothesis, because a discipline that only ever confirms its own stories has not demonstrated it has a discipline at all. For roughly a decade and a half, a substantial literature in evolutionary psychology reported that women’s preferences for cues linked to genetic quality in potential partners — facial and vocal masculinity, symmetry, dominance, testosterone-associated traits — shift systematically across the menstrual cycle, intensifying near the fertile window, particularly for short-term or extra-pair interest. The functional story was straightforward: a woman partnered for long-term investment might still benefit, in strictly genetic terms, from mating opportunistically with a genetically superior partner near ovulation, when conception is possible.

In 2014, two comprehensive meta-analyses reached opposite headline conclusions from substantially overlapping literatures. Kelly Gildersleeve, Martie Haselton, and Melissa Fales, publishing in Psychological Bulletin, pooled data translated from roughly fifty studies into a common statistical format and reported robust shifts for short-term mate assessments, though not for long-term partner evaluation, with effect sizes in the small-to-medium range [9]. In the same year, Wendy Wood, Laura Kressel, Priyanka Joshi, and Brian Louie reached the opposite conclusion from a partially overlapping set of fifty-eight independent reports. Their abstract states the finding without hedging: “fertile women did not especially desire sex in short-term relationships with men purported to be of high genetic quality (i.e., high testosterone, masculinity, dominance, symmetry). The few significant preference shifts appeared to be research artifacts. The effects declined over time in published work, were limited to studies that used broader, less precise definitions of the fertile phase, and were found only in published research” [8].

The dispute between the two teams turned out to be substantive rather than merely a matter of counting differently. The groups disagreed over whether to bundle short-term and long-term mating contexts together or analyze them separately; over how wide a “fertile window” to define around ovulation, with six-day and nine-day definitions both defended in the literature and critics arguing the wider window let researchers try multiple cutoffs until one produced significance; over how to code borderline and null results, with Gildersleeve’s team flagging seventeen effects that Wood’s coding had entered as exactly zero; and, notably, both sides acknowledged that unpublished studies in this area showed more null results than the published record did [14]. That combination — a flexible definition of the key predictor, multiple defensible analytic choices, and a publication process that favored positive findings — is close to a textbook description of how a real effect can look larger in the literature than it is, or how no effect at all can look like a real one.

The cleanest resolution came from studies designed specifically to remove that flexibility rather than argue about how to reanalyze the existing one. Benedict Jones and a large international team ran what remains one of the biggest studies of its kind: a longitudinal design tracking five hundred and eighty-four participants with repeated salivary hormone assays rather than calendar-based cycle-day estimation, testing directly whether preference for facial masculinity tracked measured hormone levels. Their own summary of the result is unambiguous: “no compelling evidence that preferences for facial masculinity were related to changes in women’s salivary steroid hormone levels,” concluding that “our results do not support the hypothesized link between women’s preferences for facial masculinity and their hormonal status” [10].

The failure mode here has a name researchers now use routinely: researcher degrees of freedom, the accumulation of individually defensible analytic choices — which window definition, which traits to test, which contexts to combine, how to handle an ambiguous cycle-day estimate — that, tried in sequence rather than fixed in advance, inflate the rate of apparent significance well above the nominal five percent threshold. Combine that with the typical sample sizes in this literature, running from the dozens to the low hundreds per study, and a publication process that selected for positive results, and the result was a research program whose apparent consensus outran what its evidence could actually support. This is not, it is worth being precise about, a case where the entire idea that hormones influence anything psychological has been shown false; narrower, hormone-assay-confirmed findings about general sexual desire persist in some studies. What specifically did not survive rigor was the headline claim that women’s preferences for masculine, symmetric, dominant-looking partners shift with fertility in the way the earlier calendar-based literature reported.

A laboratory rack of saliva-assay sample tubes on a bright bench beside a forced-choice photo-viewing booth, one tube's cycle-day label curling up half-affixed and an adjacent tube standing empty with its cap resting loose on top
Figure 6. The ovulatory-shift literature did not fail because nobody tested it; it failed because a fertile window can be defined a dozen defensible ways, and dozens of defensible choices, tried in sequence, will manufacture a significant result out of nothing.Image prompt and art direction by Brecht Corbeel; generation pending.

What to ask before believing a headline that says humans evolved to do something

Put the toolkit back together and a short, practical set of questions falls out of it, worth applying to any claim that opens with some version of “humans evolved to.” The first is whether the claim specifies which of Tinbergen’s four questions it is actually answering, and whether the evidence offered matches that question — a hormone correlation answers the mechanism question and says nothing on its own about function, phylogeny, or development, however naturally the three get run together in a headline. The second is whether the claim clears Williams’s bar for special design rather than merely borrowing its vocabulary: does the trait show a specific, structured pattern — precise, economical, aimed at a particular problem — that a general-purpose or accidental explanation would not produce, and was that pattern predicted before it was observed rather than described admiringly afterward. The third is whether universality, if invoked, is doing more work than it can bear; a trait present in every studied culture is consistent with an evolved adaptation, but it is equally consistent with every human environment reliably teaching the same lesson, and the claim needs a reason to prefer one reading over the other. The fourth is whether a phylogenetic pattern across related species, where one is available, points the same direction as the human evidence, and whether a critic’s alternative reading of that same pattern has been engaged rather than ignored. The fifth, and the one intuition gets backward most often, is what the claim assumes about heritability: high heritability is not confirmation of an ancient adaptation, and for a trait near fixation across the species, low heritability is what the adaptationist story should actually predict.

None of these five questions, on their own, can certify a hypothesis as true. Pregnancy sickness held up because it answered several of them at once, with a large and specific evidence base behind each answer, and it still carries an open question about the precise causal weight of one physiological pathway. The ovulatory-shift literature looked, for over a decade, like it had answered the special-design question convincingly, and it lost that appearance once the fertile-window definition and analytic flexibility that had generated the pattern were replaced by preregistered analysis and direct hormone measurement. The difference between those two outcomes was never the elegance of the story on offer — both stories were plausible enough to publish, cite, and repeat in the press — it was whether the story had been built to fail if the world turned out to be otherwise, and then actually checked against a world that was free to do exactly that.