Costly ornament is older than anatomically modern behavior was supposed to allow
The object that reopened this argument is a piece of red ochre roughly the size of a fist, recovered from the Middle Stone Age levels of Blombos Cave on South Africa’s southern coast. Two faces of it carry a deliberate cross-hatched design, sets of parallel incised lines crossed by a longer diagonal, made by a hand pressing a point into pigment and repeating a pattern rather than making a random mark. Henshilwood and colleagues dated the engraved pieces by thermoluminescence to a mean age of roughly 77,000 years, with an optically stimulated luminescence date of about 70,000 years on the dune sand overlying the find confirming the layer was not disturbed [1]. That is not a controversial date in the way some of what follows is controversial; it has stood for two decades as one of the more secure anchors in this literature.
Two years later the same broad team, with Karen van Niekerk and Zenobia Jacobs added, reported something arguably more informative than an engraving: forty-one small tick-shaped shells of the species Nassarius kraussianus, recovered from the same Middle Stone Age phases, each perforated in a pattern too regular to be accidental and each showing wear consistent with having been strung and worn against skin or fabric [2]. Thirty-nine came from the upper of the two relevant phases, dated to around 75,000 years. A shell bead is a more legible artifact than an engraving because its function is harder to explain away: a perforated, worn, deliberately selected shell is close to definitionally an ornament, worn on a body, meant to be seen. Before this find, the oldest known personal ornaments were about 30,000 years younger.
The African evidence for early beadwork did not stay isolated to one site. Bouzouggar and colleagues, excavating Grotte des Pigeons at Taforalt in Morocco, recovered perforated Nassarius gibbosulus shells — the same genus as Blombos, some of them stained with ochre in the same way — from layers a Bayesian model built on thirteen separate uranium-series, thermoluminescence and optically stimulated luminescence estimates placed at between 73,400 and 91,500 years old at two standard deviations, with a most probable age near 82,500 years [3]. And in 2006, Vanhaeren and colleagues pushed the practice further back still: elemental and chemical analysis of the sediment adhering to a Nassarius gibbosulus shell from Skhul in Israel matched it to a layer containing ten hominin fossils, a layer independently dated to between 100,000 and 135,000 years old — about twenty-five thousand years earlier than the Blombos beads [4]. A second shell from Oued Djebbana in Algeria, recovered far from any coastline that could have delivered it there naturally, reinforced the inference that these were carried and curated by people rather than washed in by chance.
Read together, these dates establish something that has to be taken as a fixed constraint on any theory that follows: symbolic ornamentation was not a late add-on that arrived with fully modern behavior forty thousand years ago in Europe, the old textbook story. It was present, geographically dispersed across Africa and the Levant, and apparently persistent, a hundred thousand years or more before that. Whatever explains it has to explain something old, widespread, and evidently worth the trouble of perforating a shell with stone tools before string existed to hang it from.
Divje Babe is the case where the controversy is the finding
Not every dated artifact in this literature is secure, and the honest way to handle the field is to show a case where the evidence itself remains disputed rather than only the theory built on top of it.
In 1995, excavators at Divje Babe I, a cave in Slovenia, recovered the left femur of a juvenile cave bear pierced by two clean, roughly circular holes along its shaft, from a Mousterian layer the site’s own excavators date to more than 60,000 years, a context that would put its manufacture, if it is a manufactured object, in Neanderthal hands rather than modern human ones [5]. If the holes were made deliberately, the bone would be the oldest candidate musical instrument on record by a wide margin, tens of thousands of years older than anything else discussed in this piece, and it would belong to a species that most theories below implicitly assume never made it into the picture at all.
The stakes of that “if” are easy to understate. Every adaptationist account that follows — byproduct, sexual ornament, social bonding, credible signal — is built, explicitly or not, on evidence from Homo sapiens: shell beads worn by our own lineage, cave paintings made by our own lineage, listener studies run on our own lineage today. A confirmed Neanderthal instrument would not falsify any of those accounts outright, but it would force each of them to explain why a hominin outside the modern human line, with a different social structure and a debated relationship to full syntactic language, also produced patterned sound-making objects. That is a substantially heavier evidentiary burden than any of the theories below currently attempt to carry, which is one more reason the identification question cannot be waved past.
The case for deliberate manufacture is not merely assertion. Turk and Bastiani ran two parallel experimental programs specifically to adjudicate the competing explanations. To test the carnivore hypothesis, they used metal dental casts modeled on wolf, hyena, and bear teeth to pierce twenty-nine fresh juvenile and four adult brown bear femora; the resulting perforations came out elongated, oval, or rhomboid, and reliably produced longitudinal fractures radiating from the hole, a fracture pattern the Divje Babe specimen does not show. To test the human-manufacture hypothesis, the experimental archaeologist Giuliano Bastiani used replicas of the actual Mousterian stone tools recovered from the site, in a chiseling-and-piercing technique, and produced holes with irregular, serrated edges and no conventional tool marks — a close match to the original bone, and notably a technique that leaves the kind of ambiguous edges a skeptic could otherwise point to as evidence against human work [6].
The case against is equally substantive and comes from an independent line of taphonomic work. Diedrich examined punctured cave bear femora from more than a dozen cave-bear den sites across southeastern Europe and found that hyenas, feeding on cave bear cubs, reliably produced round-to-oval punctures using their bone-crushing premolars — damage he found on roughly one-fifth of adult remains and four-fifths of cub remains at these sites — and argued the Divje Babe bone is one more instance of exactly that feeding damage, not an instrument [7]. The two camps are not disagreeing about the age of the bone or the shape of the holes; they are disagreeing about whether the same physical evidence is more consistent with a controlled Mousterian tool technique or with a hyena’s jaw, and both sides have run experiments rather than simply asserted a conclusion.
This case belongs in an honest survey of art’s origins precisely because it will not resolve into a clean number for either side of the argument that follows. If Divje Babe is a flute, the human capacity for deliberate, patterned sound production predates behaviorally modern Homo sapiens and belongs partly to Neanderthals — a fact that would matter enormously for every theory below, since it would decouple musicality from the suite of traits usually bundled with recent human cognitive evolution. If it is hyena debris, none of that follows, and the earliest secure musical instrument is tens of thousands of years younger and unambiguously modern-human. The field cannot currently tell you which world it is in, and a survey that picked a side here to make the rest of the argument tidier would be misrepresenting the state of the evidence.
What is not disputed is the next rung down in age. Excavations at Hohle Fels in southwestern Germany recovered a nearly complete flute carved from the radius of a griffon vulture, 21.8 centimeters long with a diameter of about 8 millimeters and five deliberately cut finger holes, from basal Aurignacian deposits with associated radiocarbon dates falling between 31,000 and 40,000 years before present, meaning the flute itself predates 35,000 calendar years ago; it was found close to a mammoth-ivory Venus figurine from the same research team’s excavations [8]. Nobody argues this one is a hyena’s work. Sophisticated, tuned, multi-holed wind instruments existed by then regardless of how the Divje Babe question resolves, which sets a hard floor under any theory of music’s origin: whatever produced it had fully arrived by 35,000 years ago even on the most conservative reading of the evidence.
The pictorial record tells a parallel story on a different continent. Uranium-series dating of coralloid mineral crusts that had grown over rock art at seven cave sites in the Maros karst of Sulawesi, Indonesia, established a minimum age of 39,900 years for a hand stencil and 35,400 years for a figurative animal depiction, putting representational painting in Southeast Asia on a par with, or earlier than, the oldest confidently dated cave art in Europe [9]. A later study at the same karst pushed the record further: a red-ochre painting of a Sulawesi warty pig at Leang Tedongnge, 136 by 54 centimeters, carries a minimum uranium-series age of 45,500 years, at the time of publication the oldest securely dated figurative artwork known anywhere, with a second warty-pig painting at a nearby site bracketed by dates above and below the panel to somewhere between roughly 32,000 and 73,400 years old, a wide span that the authors themselves treat as a poorly constrained minimum rather than a precise figure [10]. Cave painting, like the ochre engraving and the shell bead, is not a European invention arriving late in the story; it is old and geographically dispersed wherever anyone has looked carefully enough to date it.
Auditory cheesecake is a serious hypothesis about pleasure circuits, not an insult
Given all of that antiquity and geographic spread, the first serious theoretical claim to confront is the one that denies there is anything here for natural selection to explain at all.
Steven Pinker’s argument, made at length in How the Mind Works, is that music is what he called auditory cheesecake: “an exquisite confection crafted to tickle the sensitive spots of at least six of our mental faculties,” faculties that evolved for entirely different jobs — language, auditory scene analysis, emotional vocalization, motor control, and habitat assessment among them — and that music simply activates in combination, producing intense pleasure as a side effect rather than because music itself was selected for. Pinker’s own summary of the implication was blunt: “as far as biological cause and effect are concerned, music is useless” [11].
It is worth being precise about what this claim is and is not, because it is routinely caricatured as philistinism rather than engaged with as a hypothesis. Pinker is not saying music does not matter to people, does not move them, or is not worth making. He is making a narrower, testable evolutionary claim: that the capacity to be moved by music is a byproduct of machinery built for other purposes, in the same category as the human ability to enjoy a sweet, high-fat dessert that never existed in the ancestral environment in the calorie-dense, effortlessly available form a bakery now provides. Cheesecake exploits a taste system built to seek out ripe fruit and animal fat; on this account music exploits language-processing circuitry built to track pitch contour and rhythm for speech, motor circuitry built for coordinated movement, and emotional-call circuitry built to read vocal affect in others. The byproduct account does not require any of those component capacities to be small or the pleasure they produce to be trivial — it requires only that music itself, as opposed to the faculties it borrows, never had to pay its own way in survival or reproduction.
This is a coherent, falsifiable position, and its predictions are distinct from a straightforward adaptationist account: a byproduct should show no dedicated neural architecture of its own, should be dissociable from the faculties it draws on only in the sense that damage to those faculties should degrade musical response correspondingly, and should not need to appear reliably in development or across cultures in any particular form, because nothing was selected to guarantee its appearance. Later sections return repeatedly to evidence that bears on exactly these predictions — universality of form, developmental timing, and whether musicality dissociates from language and motor ability in principled ways — because the byproduct hypothesis, unlike a dismissal, actually generates things to look for.
The ornament logic predicts a mating-effort signature the data only partly deliver
The competing account with the deepest intellectual pedigree treats art and music not as leftovers from other systems but as products of sexual selection in their own right — ornaments in the technical sense Darwin gave the word, costly traits that exist because they influenced who mated with whom.
Darwin himself set out the logic for music specifically, and did so before there was any archaeological evidence at all to test it against. In The Descent of Man, he wrote that “some early progenitor of man probably first used his voice in producing true musical cadences, that is in singing, as do some of the gibbon-apes at the present day; and we may conclude from a widely-spread analogy, that this power would have been especially exerted during the courtship of the sexes,—would have expressed various emotions, such as love, jealousy, triumph,—and would have served as a challenge to rivals” [15]. The claim is specific: musical vocalization predates articulate language, and its original job was courtship display and rivalry signaling, the same job birdsong does in species where only or mostly males sing to attract mates and warn off competitors.
Geoffrey Miller’s The Mating Mind, and the chapter-length argument he published the same year specifically on music, generalized Darwin’s logic to the whole of human creative behavior. Miller’s claim is that art, music, wit, and elaborate language are best understood as fitness indicators in Amotz Zahavi’s sense: costly, hard-to-fake displays of underlying qualities — cognitive ability, fine motor control, physical stamina, genetic quality — that a potential mate cannot observe directly but can infer from an ornament expensive enough that only a genuinely high-quality individual could afford to produce it. As direct empirical support, Miller pointed to demographic patterns in recorded output: across jazz, rock, and classical album catalogs, male musicians produced roughly ten times the recorded output of female musicians, with output peaking around age thirty — the same age at which mating effort, by several independent measures, peaks in men. Miller was careful to flag the limits of this evidence himself, noting the sample was restricted to musicians successful enough to appear in discographies and encyclopedias, and explicitly called for more direct quantitative work linking musical or artistic ability to mate preference and to heritable fitness measures such as aerobic capacity and general intelligence [12].
That direct work has been done in the two decades since, and the honest summary is that it supports Miller’s theory only partially, and only for one sex. Clegg, Nettle, and Miell surveyed 236 visual artists and modeled the relationship between artistic success — a composite of exhibition history, sales, and self-reported status — and lifetime sexual partner count. For the eighty-five male artists in the sample, artistic status was the single significant predictor of mating success, explaining thirty-five percent of the variance in partner count; for the hundred and fifty-one female artists, the same model found no such relationship, and the strongest predictor of their outcomes instead was relationship length, in the opposite direction, suggesting a strategy oriented toward mate quality rather than mate number [13]. This is exactly the sex-differentiated pattern a naive reading of sexual selection theory would predict for a species with asymmetric parental investment, and it is real evidence in Miller’s favor — but it is evidence for male creative display specifically, not for a general theory of why humans of both sexes make art.
A more recent test complicates the picture further by leaving the WEIRD populations — Western, educated, industrialized, rich, and democratic — in which almost all of this research had previously been conducted. Lebuda and colleagues tested a hundred and thirty-three members of the Meru people in Kenya, a population without modern contraception and with natural fertility, using a figural creativity test and self-reported numbers of spouses and children. Creative potential negatively predicted number of offspring, and that relationship was fully mediated by number of spouses: more creative individuals in this sample more often stayed single and had no children at all [14]. That is close to the opposite of the pattern Miller’s theory would predict if creativity functioned as a universal, cross-cultural mating advantage rather than a signal whose payoff depends heavily on the particular mating market — monogamous, high-status, urban, filled with strangers to impress — in which most of the supporting data happen to have been collected. Taken together, the direct evidence audit does not refute sexual selection as a factor in art and music; it shows the effect, where it exists, is smaller, more sex-specific, and more culturally contingent than the ornament analogy by itself would lead anyone to expect.
Song carries a cross-cultural grammar of function, and two rival stories about why
A different kind of evidence entirely comes from asking not whether artistic ability predicts mating success, but whether the forms music takes are themselves universal in ways that would need explaining regardless of who is doing the singing or why.
Mehr and a large international team built two independent datasets to test this directly: an ethnographic text corpus of 4,709 descriptions of song drawn from sixty societies sampled to represent the world’s cultures, and a discography of 118 field recordings from eighty-six societies across thirty geographic regions, each recording classified into one of four functional categories — dance songs, lullabies, healing songs, and love songs. They then had roughly 29,357 online listeners from around the world, none of them ethnomusicologists, classify unfamiliar recordings by function alone, from sound with no translated lyrics or cultural context, generating 185,832 individual ratings. Listeners correctly identified a song’s function 42.4 percent of the time against a 25 percent chance baseline, with dance songs the easiest to identify at 54.4 percent accuracy and love songs the hardest at 26.2 percent — barely above chance, and the category where cross-cultural agreement about what love song even sounds like appears weakest. A machine classifier trained on musical features rather than listener intuition reached 50.8 percent. Across both corpora, musical behavior varied along three recurring dimensions — formality, arousal, and religiosity — that together explained roughly a quarter of total variance, but variation within a single society consistently exceeded variation between societies by a factor of about six [16].
This is a genuinely important and well-powered result: it shows real, detectable, cross-culturally shared regularities linking a song’s form to its social function, at a scale and with a design few studies in this literature can match. What it does not do, on its own, is settle why those regularities exist, and two rival evolutionary accounts were published side by side in the same 2021 volume of Behavioral and Brain Sciences, making this one of the few places in the field where the live disagreement is visible in a single citation rather than scattered across decades of literature.
Savage and six co-authors argued that social bonding is the unifying function beneath music’s evolution — not a rival to mate-selection, parental-care, or coalition-signaling accounts, but the deeper mechanism that explains why each of those specific contexts recruits music at all. Their claim is that musicality let hominins extend social bonding to group sizes larger than grooming, the bonding mechanism available to other primates, could support, through a feedback loop of production, perception, prediction and reward built on rhythmic synchronization and pitch harmonization, with the whole system shaped by gene-culture coevolution as musical behaviors that first spread as cultural inventions gradually acquired genetic scaffolding [17].
Mehr, Krasnow, Bryant, and Hagen, publishing the companion target article in the same issue, rejected the social-bonding synthesis as too diffuse to generate sharp predictions and proposed instead that music is a credible signal operating in two specific, testable contexts. First, coordinated rhythmic group displays credibly signal coalition size, strength, and coordination ability to rivals and allies alike, because faking synchronized rhythm at scale is hard to do without the coalition actually possessing the coordination it advertises. Second, infant-directed song credibly signals a caregiver’s attention and commitment to an altricial infant who cannot otherwise verify it, because the fine, real-time responsiveness a parent’s singing shows to an infant’s shifting affective state is a costly, hard-to-fake cue of sustained attention in a way that generic vocalization is not [18]. Both papers accept the same universality data Mehr’s own 2019 study established; they disagree, sharply and explicitly, about what kind of selective pressure produced it, and neither side’s commentary in that BBS issue treats the matter as settled.
Infants arrive tuned for music, and no other animal is tuned quite the same way
Two further lines of evidence bear directly on the theories above without deciding between them, and both come from looking at when and in whom musical responsiveness appears.
Sandra Trehub’s research program, spanning several decades at the University of Toronto, established that human infants arrive with musical perceptual abilities well beyond what exposure alone would predict, including a documented capacity to acquire complex, non-native rhythmic structures with a facility that adults, in her comparative work with Erin Hannon, do not retain — evidence for an early, broad perceptual window for musical structure that appears to narrow with cultural exposure rather than one that has to be built up gradually from nothing [19]. Trehub’s own studies also found that singing to infants keeps them settled for roughly twice as long as speaking to them does, a finding replicated widely enough that infant-directed song is now treated as a distinct, cross-culturally recognizable register of vocal behavior in its own right rather than merely speech set to a tune. This developmental evidence is compatible with several of the theories above rather than adjudicating cleanly among them: an innate, early-arriving perceptual bias toward musical structure is exactly what a genuine adaptation should show, but it is also compatible with a byproduct account if the perceptual machinery being recruited — auditory scene analysis, prosody tracking — is itself present from birth for unrelated linguistic reasons.
The comparative evidence cuts in the opposite direction from what a naive analogy to birdsong would suggest, and the disanalogies matter more than the surface similarity Darwin originally drew on. Birdsong, in most species that have it, is typically restricted to males, tied tightly to a breeding season, and functions almost exclusively in territorial defense and mate attraction — a genuinely narrow, single-purpose signal. Human music is produced and enjoyed by both sexes and every age class, appears in contexts with no plausible courtship function at all (work songs, funeral dirges, solitary humming, lullabies sung by mothers to infants who cannot yet choose a mate), and coexists with a fully developed language system that birds, so far as anyone can tell, do not have. Miller’s own analogy to bird song captures something real about costly vocal display, but the mismatch in when, by whom, and in what contexts the display actually occurs is exactly the kind of disconfirming detail a sexual-selection account has to explain away rather than simply point past.
Three accounts remain live, and the discriminating test has not been run
None of the three positions surveyed here — byproduct exploitation of pleasure circuits, sexually selected fitness display, and coevolved social-bonding or credible-signaling systems — has been eliminated by the evidence that currently exists, and a survey that declared a winner would be overstating what any single study, including the best-designed ones, actually shows.
Each account does make at least one prediction the others do not share, which is what keeps this a live scientific dispute rather than an unfalsifiable shouting match. The byproduct account predicts musicality should show no dedicated developmental trajectory independent of language and motor development, and should be dissociable only in step with damage to the systems it borrows from; evidence of a distinct developmental timetable for musical perception that outruns general language development, or of amusia that leaves language and motor control intact, would weigh against it. The sexual-selection account predicts a mating-effort signature — greater investment in creative display around peak reproductive age, concentrated more heavily in one sex, correlating with actual mating outcomes across diverse mating systems, not only industrialized, monogamous ones; the Lebuda result in a natural-fertility population is exactly the kind of finding that should worry defenders of the strong version of this account, and a wider set of comparable non-WEIRD replications, in either direction, would move the needle substantially. The credible-signaling account predicts infant-directed song specifically, rather than music in general, should show the sharpest, most consistent cross-cultural regularities, since it is tied to a narrow, high-stakes function that cannot easily drift; Mehr’s own 2019 data, in which lullabies were not the most easily identified category, is a data point in tension with the strongest form of this prediction rather than confirmation of it. The social-bonding account predicts that group musical synchrony specifically, independent of individual display or infant care, should carry measurable effects on cooperation and coalition cohesion that scale with group size in the way grooming does not; that specific comparative claim remains less directly tested against alternatives than the theory’s proponents’ confidence would suggest.
What would actually discriminate among these accounts is a research program none of the existing studies fully delivers: longitudinal developmental data tracking musical perception against language and motor milestones across many cultures at once; mating-outcome studies in non-WEIRD, non-monogamous societies large enough to detect sex-specific effects reliably; and direct physiological comparisons of coalition-signaling contexts against infant-care contexts within the same populations, rather than theoretical arguments about which context better fits the existing correlational data. Until that program exists, the responsible summary of three decades of evolutionary argument about art is not a verdict. It is that art and music are old enough, costly enough, and universal enough that something evolutionary is almost certainly going on, and specific enough in their disputed details — a hyena’s jaw or a Neanderthal’s chisel, a mating display or a parent’s attention, a spandrel or an adaptation — that nobody yet gets to say which something it is.