Ask a laboratory psychologist, a field ethnographer, and an archaeologist how cumulative culture works, and you will get three different research programs before you get three different answers. Each has built its own instrument for a phenomenon that leaves no single kind of trace: humans build on the achievements of others across generations, in a way that appears to have no close parallel in the rest of the animal world outside a handful of contested cases [1]. Explaining that pattern — why it happens, what mechanisms sustain it, and why it sometimes stalls or reverses — has never been the property of one discipline, and the differences between the three main approaches are not a matter of taste. They trade specific, nameable advantages for specific, nameable costs, and the trade is different in each case.
This article compares the laboratory transmission-chain experiment, cross-cultural ethnographic fieldwork, and archaeological analysis of durable material culture on the dimensions that actually distinguish them: how much the researcher controls, how realistic the setting is, how deep in time the method can reach, and what kinds of claims about imitation, teaching, language, population structure, innovation, loss, transmission fidelity, and cooperation each one can support. None of the three is a lesser version of another. Each answers a question the other two cannot.
What “cumulative culture” means, and why the definition already picks a method
Cumulative culture, sometimes called the “ratchet effect,” refers to a population’s ability to build on previous innovations across generations so that culturally transmitted practices become more complex or effective over time than any single individual could invent alone [1]. A widely used review breaks the concept into components that different methods are suited to measuring separately: fidelity of transmission (how accurately information passes from one learner to the next), innovation (the introduction of new variants), and selection or retention (which variants persist and which are lost) [8].
That decomposition matters here because each research approach was built to see one or two of those components clearly at the cost of the others. A transmission chain in a laboratory can measure fidelity and innovation with near-total precision because the researcher constructed the chain. An ethnographer watching a real apprenticeship can see how teaching, correction, and social pressure actually operate, but cannot rerun the apprenticeship with one variable changed. An archaeologist looking at ten thousand years of stone tools can see whether a technology’s complexity trended upward, downward, or stalled, but cannot observe a single act of imitation that produced any of it. The three methods are not converging on the same evidence from different angles; they are collecting different kinds of evidence entirely.
The laboratory transmission chain: total control, a manufactured world
The core method, formalized by Christine Caldwell and Ailsa Millen, seats a chain of participants where each person attempts a task — building the tallest possible tower from sticks and clay, folding the farthest-flying paper airplane — after seeing only the output of the person before them, not the earlier steps that produced it [2]. Performance climbs across chain “generations” in a way that individual, isolated attempts do not reliably match, which is taken as direct behavioral evidence that transmission across a chain of learners is doing real work, separate from each person’s own trial-and-error [2].
The appeal of this design is causal control that neither of the other two approaches can offer. A researcher can hold the task fixed, vary exactly one thing — whether learners can imitate the process or only see the product, whether the group is fully connected or partitioned into subgroups, how many people occupy each chain link — and attribute any resulting difference in outcome to that one manipulated variable. Maxime Derex and Robert Boyd used exactly this kind of design to show that groups whose members are only partially connected to one another, rather than uniformly interacting with everyone, can sustain more diverse solutions and reach more complex outcomes than fully connected groups of the same size, because partial isolation lets alternative variants survive long enough to be improved rather than converging early on one adequate solution [3]. That is a genuinely causal claim about population structure, and it is only available because the researcher built and controlled the population structure directly.
The laboratory approach is equally suited to separating imitation from other social-learning mechanisms — copying the specific actions someone used versus copying only the result they achieved, sometimes called emulation — because the researcher can restrict what information a learner is allowed to see and then compare outcomes [1]. This distinction sits near the center of debates about why cumulative culture appears to be far more developed in humans than in other apes: comparative work summarized in the same literature suggests that non-human great apes learn socially in ways that support the transmission of behavior but rarely support the accurate copying of others’ actual technique, which caps how far their traditions can be improved upon by later individuals [1].
What the laboratory buys in control, it spends in realism, and this is not a minor caveat. A chain of undergraduates building a tower from spaghetti and clay over forty-five minutes shares little with a Kalahari toolmaker’s decade of apprenticeship: the tasks are arbitrary, the stakes are trivial, the timeframe is compressed from years to an afternoon, and the participants are drawn overwhelmingly from a narrow demographic slice of industrialized, educated populations. Whether a mechanism demonstrated in that setting is the same mechanism operating in a real subsistence economy is an inference, not a direct observation. The laboratory experiment is the strongest method available for establishing that a mechanism can produce cumulative improvement under controlled conditions; it is a much weaker method for establishing that the same mechanism is the one actually operating, or the dominant one, in any particular human society.
Cross-cultural ethnographic fieldwork: ecological validity, weak causal leverage
Where the laboratory manufactures a task and a population, ethnographic fieldwork records a task and a population that already exist. Barry Hewlett and L. Luca Cavalli-Sforza’s study of the Aka, a Central African forager population, quantified who taught whom, at what age, and for which specific skills, and found that basic subsistence skills were typically acquired very early, largely from parents, while specialized knowledge — including some medicinal and ritual knowledge — continued to be transmitted well into adolescence and adulthood, often from individuals outside the immediate family [5]. That level of demographic and social specificity about how transmission actually distributes through age and kinship in one real community has no laboratory equivalent, because no laboratory chain lasts long enough, or involves real kin relationships, real economic stakes, or real cultural institutions of teaching.
This is the method’s central strength: ecological validity. When an ethnographer documents that some skills pass vertically from parent to child while others pass obliquely from unrelated specialists, that is a fact about a real transmission system, not an inference from a proxy task. It bears directly on questions the laboratory struggles to reach at all — what role explicit teaching plays relative to unsupervised observation and imitation, how transmission pathways differ by the type of skill being taught, and how cooperation and division of labor structure who has access to which knowledge.
The cost is that ethnography offers very limited causal leverage. A researcher observing the Aka cannot experimentally remove teaching from one subgroup and compare it with an otherwise identical subgroup that retains teaching; the myriad other differences between any two real communities — subsistence base, group size, contact history, environment — confound any comparison a fieldworker draws across societies. Findings are also necessarily particular to the community studied at the time it was studied, and generalizing from one forager population’s transmission pattern to “human cultural transmission” in general requires either replication across many independent field sites or explicit theoretical argument about why the pattern should generalize — neither of which the original observation itself supplies. Comparative studies of technological complexity across island societies, discussed below, address generalization by assembling many cases rather than one, but that move trades some of ethnography’s fine-grained detail for breadth.
Archaeological material analysis: the only method with real time depth
Neither the laboratory nor ethnographic fieldwork can observe a process that unfolds over thousands of generations. Archaeology can, because durable material culture — stone tools above all — survives long after the people who made it, and can be dated, measured, and compared across an assemblage spanning enormous timescales. A 2024 analysis of stone tool complexity spanning roughly 3.3 million years argued that measurable complexity in stone tool technology remained comparatively flat for most of that span and only began a sustained increase during the Middle Pleistocene, which the authors interpret as evidence for when cumulative cultural processes, as opposed to simple technological continuity, became a significant force in hominin technology [9]. No laboratory chain or ethnographic study could speak to a question posed at that timescale; the evidence physically does not exist in any other form.
Stephen Lycett’s work on cultural phylogenetics extends this by treating variation among stone tool assemblages the way biologists treat variation among species: discretizing tool shape and technological attributes into comparable traits and using quantitative methods borrowed from evolutionary biology to reconstruct which assemblages likely descended from which, and how transmission fidelity and drift shaped the resulting pattern of similarity and difference across sites and periods [7]. This lets archaeologists ask questions about transmission fidelity and population-level descent using nothing but the durable objects themselves, without ever having observed a single teaching event.
The dimension where archaeology genuinely rivals laboratory experiments, rather than merely complementing ethnography, is population size and structure. Michelle Kline and Robert Boyd’s study of marine foraging tool kits across ten Oceanic island societies found that islands with smaller, more isolated populations tended to have less complex tool repertoires than islands with larger populations or more contact with neighboring groups, which the authors interpret as consistent with models in which a larger pool of individuals sustains more innovation and reduces the rate at which useful skills are lost to chance [4]. This is a real-world, historically grounded test of a claim about population structure and cumulative complexity — the same claim Derex and Boyd tested causally inside a laboratory chain [3] — arrived at through comparison across naturally occurring populations rather than experimental manipulation.
But archaeology’s reach into deep time is bought at a specific cost: it can describe patterns of loss and gain in material complexity with high confidence, while saying comparatively little about the underlying social-cognitive mechanism that produced them. A stone tool assemblage cannot tell you whether the people making it learned primarily by imitation, teaching, or independent trial and error, because the physical remains do not preserve the learning event itself, only its downstream product — and multiple different learning mechanisms can, in principle, produce similar patterns of technological continuity or change. Where the transmission-chain experiment isolates mechanism at the cost of realism, and ethnography records mechanism embedded in reality at the cost of causal control, archaeology records outcome at a scale no other method reaches, at the cost of direct access to mechanism at all.
Chimpanzee culture: a case where all three methods have collided
The chimpanzee tool-use literature is an unusually clear place to see the three approaches applied to the same broad question and reaching different conclusions from different evidence. Field observations of chimpanzee populations across Africa documented dozens of behavioral patterns — tool use, grooming styles, courtship displays — that vary between populations in ways not explained by genetics or local ecology, and the researchers who compiled this record explicitly labeled the pattern “cultures in chimpanzees” [6]. That is an ethnographic-style field finding: real populations, real behavioral variation, documented as it exists.
Laboratory transmission-chain work with captive apes has been used to ask a narrower, mechanistic question this field record cannot answer on its own: whether the observed behavioral variation could plausibly be sustained by the same high-fidelity imitation and cumulative improvement seen in human chains, or whether it more likely reflects each individual rediscovering a behavior under similar ecological pressure, assisted only loosely by watching others [1]. The laboratory evidence on ape social learning generally finds copying of outcomes more readily than copying of the precise technique used to reach them, which constrains how much any single ape tradition can be incrementally improved across generations compared with human traditions [1]. Archaeological methods have, separately, been applied to hominin and even chimpanzee stone tool use to ask how far back in time comparable material traditions extend and whether their complexity trajectories resemble the pattern found in the human lineage [9]. No one of these three literatures alone settles whether chimpanzee behavioral variation constitutes “culture” in the cumulative sense central to the human case; the field record establishes that variation exists, the laboratory record constrains what mechanism could sustain it, and neither speaks to material continuity over the timescales archaeology addresses for the human lineage. The disagreement in this literature about how “chimpanzee culture” should be characterized is a genuine, unresolved one among researchers who study it, not a gap this article can close, and it is worth naming as such rather than adjudicating.
Comparing the three approaches directly
Laid side by side, the trade-offs are structural rather than incidental to any one study’s design.
Ecological validity. Ethnographic fieldwork is highest: the community, the skills, and the stakes are real. Archaeology is intermediate: the material record is real, but the social process that produced it must be inferred. The laboratory is lowest by design: tasks are arbitrary and timeframes are compressed, which is precisely what allows control.
Causal control. The laboratory is highest: a researcher can manipulate population connectivity, chain length, or the information available to each learner, one variable at a time [3]. Ethnography offers almost none, relying instead on natural variation and researcher interpretation. Archaeology offers a different kind of leverage — comparison across many naturally varying populations, as in the Oceanic tool kit study — which is stronger than ethnographic single-case observation but still short of true experimental manipulation [4].
Time depth. Archaeology is unmatched, reaching into millions of years where fossilized and durable materials survive [9]. Ethnography is bounded by a researcher’s career and the historical record available for a living population. The laboratory is bounded by a single study’s duration, typically hours to at most a few months of chained sessions.
Access to mechanism. Here the order reverses. The laboratory offers the cleanest access to mechanism because the researcher defines what information can pass between learners [2]. Ethnography offers real but confounded access, since teaching, imitation, and independent practice are all present simultaneously in any observed skill transfer [5]. Archaeology offers the least direct access to mechanism of the three, since the physical record preserves outcomes but not the cognitive or social process by which they were reached [7].
Generalizability. None of the three methods generalizes for free. A laboratory result generalizes only as far as the task and population sampled resemble the target case, which is frequently a serious limitation given how heavily the experimental literature draws on university-adjacent populations. An ethnographic finding generalizes only as far as the studied community resembles others, and requires either replication or explicit comparative work to extend, as in the Oceanic island comparison [4]. An archaeological pattern generalizes only as far as the preserved material record represents the full range of past behavior, given that only durable materials like stone survive at all, systematically under-representing organic technologies such as basketry, cordage, or wooden tools that likely also carried cumulative traditions but rarely fossilize.
What none of the three approaches can settle alone
It is tempting, once the trade-offs are laid out this plainly, to ask which approach is “best” for studying cumulative culture. That question has no defensible answer, because the three methods are not competing measurements of the same quantity; they are answers to different questions that happen to share a name. Whether imitation or emulation dominates a given ape or human learning task is a question suited to laboratory manipulation. Whether teaching, in the sense of one individual actively structuring another’s learning, is common or rare in a given real society, and for which skills, is a question suited to ethnographic fieldwork. Whether technological complexity in a lineage rose, fell, or stalled across hundreds of thousands of years is a question only the archaeological record can address at all.
Language sits awkwardly across all three. Laboratory chains can test whether adding a verbal-instruction condition to a physical task changes fidelity or speed of improvement, holding the task itself constant [2]. Ethnographers can observe how much of real teaching, such as among the Aka, is verbal versus purely observational and how that balance shifts with the type of skill and the age of the learner [5]. Archaeology cannot observe language directly at all; claims about language’s role in the deep history of cumulative technological change are necessarily inferential, built from indirect proxies such as the complexity or standardization of an assemblage, and should be read as more speculative than either the laboratory or ethnographic evidence discussed above.
Two forecasts follow from this comparison, both conditional and both falsifiable rather than confident predictions. First, a horizon of roughly the next decade: if large-scale, long-duration transmission-chain experiments — chains run over months rather than hours, with real skill acquisition rather than arbitrary tasks — become more common as recruitment and remote-participation infrastructure improve, the gap between laboratory ecological validity and ethnographic ecological validity should narrow measurably, visible as an increase in published chain studies whose tasks require sustained practice rather than single-session assembly. If instead chain studies remain dominated by short single-session tasks a decade from now, that would disconfirm the forecast. Second, on the same horizon: if archaeological methods for inferring social-learning mechanism from material remains — for example, statistical signatures of copying fidelity in tool-shape variation, building on approaches like Lycett’s cultural phylogenetics [7] — continue to mature, expect more explicit cross-references between archaeological and laboratory literatures attempting to validate one method’s inferences against the other’s controlled findings; continued near-total separation between these two literatures a decade from now, with archaeology papers rarely citing laboratory mechanism studies or vice versa, would disconfirm this forecast.
None of this licenses treating the three approaches as substitutable, and none of it licenses ranking them. A finding from a transmission-chain experiment about population connectivity is not more or less true than a finding from an Oceanic tool-kit comparison about population size; they are answers to related but distinct questions, obtained by instruments built for different jobs, and the honest synthesis of cumulative-culture research depends on keeping that distinction explicit rather than collapsing it into a single ranked league table of methods.