Three research communities dig, sequence, and code their way toward the same question — how did small farming settlements turn into cities, states, and empires — and for most of the twentieth century they barely spoke to each other. A dirt archaeologist stratifying a mudbrick mound in a single trench, a population geneticist extracting DNA from a petrous bone in a clean room, and a historical sociologist coding five hundred polities into a shared database are not doing the same kind of science, even when they are describing the same transition in the same place. This article puts the three approaches side by side on the questions that matter most for early complex societies — domestication, settlement growth, disease, hierarchy, trade, infrastructure, state formation, and collapse — and asks what each one can and cannot see. None of the three is a substitute for the others; the interesting findings of the last fifteen years have mostly come from forcing them to answer to each other.
Three ways of asking the same question
Settlement archaeology and excavation reads a site the way a geologist reads a cliff face: layer by layer, in strict superposition, with each layer datable relative to the ones above and below it and, with radiocarbon or other absolute methods, in calendar years. Its unit of analysis is the single site or, at best, a regional settlement pattern built from surveying many sites. Its evidentiary strength is resolution — a trench through a tell can show a mudbrick wall built, burned, and rebuilt three times inside a few centuries, with grain stores, hearths, and burials in direct stratigraphic relationship to each other. Its weakness is scale: even a well-excavated site is one case, and generalizing from one case to “how cities formed” is an inference the data cannot fully support on its own.
Ancient-DNA population genomics reads a different archive: not the buried structure but the buried body. Sequencing nuclear and mitochondrial DNA from bone or tooth samples recovers ancestry, population movement, admixture between groups, kinship among individuals in a cemetery, and the genomes of ancient pathogens preserved incidentally in the same tissue. Its strength is that population-scale processes — migration, replacement, admixture, epidemic history — leave a genetic signature that stone tools and pottery styles cannot show directly, because artifact styles can spread by imitation without any accompanying movement of people, while ancestry cannot. Its weakness is that DNA says nothing about law, ritual, taxation, or which political institution held power; a genome cannot tell you whether a graveyard belonged to free farmers or hereditary elites unless combined with grave goods, isotopic diet data, or other contextual evidence.
Comparative historical databases, of which Seshat: Global History Databank is the most developed example, take the opposite trade-off from excavation. Rather than going deep on one site, Seshat’s research team — co-founded by Peter Turchin, Harvey Whitehouse, and Pieter François — codes hundreds of past polities across ten thousand years of history and prehistory on a shared set of variables: population size, territory, administrative levels, information systems, religious characteristics, and more, drawing on the published work of historians and archaeologists rather than new fieldwork [1] [2]. Its strength is statistical power: with enough independently coded cases, a hypothesis like “does state formation precede or follow moralizing religion” becomes a testable cross-cultural pattern rather than a single narrative. Its weakness, made vivid by its own history, is that the entire method depends on the quality, consistency, and genuine independence of every underlying case coding — and coding is itself an interpretive act performed by people reading secondary literature, not a direct measurement.
Domestication: grain morphology versus genomic selection signatures
Archaeobotanists identify domestication in the field and the sorting tray, not in a genome browser. The classic marker is the shift from a shattering seed rachis, which lets wild cereals disperse their own seeds, to a tough, non-shattering rachis that keeps the grain on the stalk for harvesting — a trait with no advantage to the wild plant and every advantage to a human collecting a crop by hand. Measured across thousands of charred grains recovered by flotation from Near Eastern sites, the frequency of the domesticated rachis form climbs over centuries, giving archaeology a slow, spatially explicit timeline of when and where domestication traits become common in the archaeobotanical record. This is a strength few other methods share: fine chronological resolution tied to a specific site’s stratigraphy.
Ancient-DNA and modern comparative genomics contribute a different piece: they can identify the specific genes under selection during domestication and, in living crop populations, estimate how many independent domestication events occurred and where their wild ancestors originated. But sequencing ancient plant DNA at scale is far harder than sequencing ancient human or animal DNA — plant tissue degrades faster and contains more inhibitors — so most of what is known about the genetic basis of crop domestication still comes from comparing modern crop and wild-relative genomes rather than dense ancient plant genome time series. The two approaches are complementary rather than substitutable: archaeobotany dates the trait’s spread on the ground; genomics identifies the molecular mechanism and lineage. Neither can be inferred safely from the other alone.
Settlement growth and the case against agriculture as origin of complexity
One of the clearest recent overturns in this field came directly from excavation, not from genetics or comparative coding. At Göbekli Tepe in southeastern Turkey, Klaus Schmidt’s team, working since 1994, excavated monumental T-shaped limestone pillars, some reaching over five meters and weighing many tons, carved with animal reliefs and arranged in circular enclosures — construction Schmidt dated to roughly 9600 BCE, which is before any evidence of domesticated plants or animals at the site [9]. This is a fact established by excavation and radiocarbon dating specifically: without a stratified, datable sequence at one site, there would be no way to establish that monumental, apparently ritual construction on this scale preceded agriculture in this region rather than following it. The finding reversed a long-standing assumption — that food surplus from farming was a precondition for large-scale communal construction — and it is a finding of a kind that neither ancient DNA nor a comparative database was positioned to make, because it required recognizing an anomaly at a single, specific place before it could become a testable general claim.
Settlement archaeology’s regional-survey variant then supplies what single-site excavation cannot: the scaling pattern of settlements across a landscape over time, showing whether population aggregated gradually into fewer, larger sites or whether early cities appeared abruptly. This is analysis rather than raw fact — inferring “gradual aggregation” from a settlement-size distribution requires assumptions about survey coverage and site visibility that vary by region and by how much later occupation or erosion has disturbed earlier layers.
Disease: what only ancient DNA can show directly
Disease history is the domain where ancient DNA has produced results that excavation alone structurally cannot, because pathogens rarely leave macroscopic skeletal traces before killing their host, and even where they do (tuberculosis, treponemal disease), skeletal lesions cannot distinguish related pathogen strains or establish a phylogeny. Reconstructed ancient Yersinia pestis genomes from Bronze Age burials across Eurasia, dated to roughly 5,000 to 2,500 years before present, showed that plague-causing bacteria were circulating far earlier and more widely than any documentary or skeletal evidence had suggested, and that at least two genetically distinct lineages existed in parallel across a huge geographic range [7]. A separate genome, reconstructed from an approximately 3,800-year-old individual in the Samara region of Russia, carried the ymt virulence gene associated with flea-borne transmission — the trait that enables bubonic plague specifically — showing that this transmission capability evolved more than a thousand years earlier than researchers had previously proposed based on later, better-attested pandemics [8]. Both are fact-level findings: specific genomes, specific mutations, specific ages, established by direct sequencing.
What ancient DNA on its own cannot establish is societal consequence — whether these Bronze Age plague lineages caused population decline, contributed to the disappearance of Late Neolithic cultures in parts of Europe, or spread through trade networks versus other contact routes. Those are analytical claims that require combining the genomic dating with independently derived archaeological population-proxy data (radiocarbon-dated site counts, settlement abandonment sequences) and, ideally, comparative data on other regions that did not experience population decline over the same interval, to check whether the timing correlation could be coincidental. This is a case where all three approaches genuinely need each other: DNA supplies the pathogen’s existence and age, excavation supplies population-proxy trends, and a comparative framework is needed to judge whether the correlation between the two is stronger than chance.
Hierarchy, trade, and infrastructure: converging lines, different blind spots
Social hierarchy inside early settlements is read mainly from mortuary archaeology and settlement layout: differences in grave goods, house size, and access to storage between contemporaneous burials or households in the same excavated horizon. This is a strong, direct signal precisely because it is measured within one stratigraphic context, holding time and place constant — it is a comparison archaeology is structurally well suited to make. Ancient DNA adds a layer current mortuary archaeology alone cannot: kinship. Sequencing multiple individuals from one cemetery can establish whether high-status burials cluster within a single biological lineage, which speaks directly to whether early hierarchy was becoming hereditary — a question about social structure that skeletal grave-good analysis alone can suggest but not confirm, since two unrelated individuals can be buried with similar goods for reasons unrelated to descent.
Trade and exchange networks are established primarily through sourcing studies — chemically or isotopically matching an artifact found at one site to a raw-material source elsewhere, most classically for obsidian, whose trace-element signature can be matched to a specific volcanic outcrop using portable X-ray fluorescence or laser-ablation mass spectrometry. This produces hard, falsifiable point-to-point evidence of exchange distance and direction that neither ancient DNA nor comparative databases can substitute for; a genetic similarity between two populations does not establish that they traded, and a comparative database’s “trade” variable is a coded summary of exactly this kind of sourcing evidence collected by others, not an independent measurement.
Infrastructure — irrigation canals, city walls, granaries, road networks — is almost entirely the domain of excavation and remote sensing (aerial and satellite imagery, LiDAR), because infrastructure is by definition a physical, spatially extended construction that leaves traces in the ground or on the surface. A comparative database can code “presence of irrigation” or “settlement hierarchy levels” as a variable, but that coded value is downstream of, and only as reliable as, the underlying excavation and survey reports it summarizes.
State formation: the comparative database’s home turf, and its cautionary tale
State formation — why some regions consolidated into centralized polities with standing bureaucracies, taxation, and law while others did not — is the question comparative databases are built to answer, because it requires comparing outcomes across many independent cases holding for confounding variables, which is exactly the kind of test a single excavated site cannot run. Seshat now holds roughly 300,000 coded records covering about 500 past societies over ten thousand years [2], drawn from published historical and archaeological scholarship rather than new fieldwork, which lets researchers ask statistical questions like whether increases in social scale statistically precede or follow the appearance of moralizing “big gods” doctrines across dozens of independent world regions.
The most important thing to say about this method is also a cautionary tale, and it is worth stating plainly rather than glossing over. In 2019, a Seshat-based study concluded from that same database that complex societies with populations above roughly one million people tended to precede, rather than follow, the appearance of moralizing high gods, published in Nature under the title “Complex societies precede moralizing gods throughout world history” [3]. The paper was retracted after independent researchers identified problems with how some cases had been coded and with the statistical independence of some data points across the sample. A subsequent reanalysis by an overlapping set of authors, published in Religion, Brain & Behavior as a “review and retake,” revisited the corrected data and reported that the broad pattern — complexity increases preceding, rather than following, the largest jumps in moralizing religious doctrine — held up under the corrected coding, though the original headline framing had overstated the strength and universality of the claim [4]. This episode is the clearest illustration available of comparative history’s central vulnerability: the method’s statistical power is entirely contingent on the quality and true independence of hundreds of individual case codings performed by people interpreting secondary sources, and an error or non-independence buried in that coding process can propagate into a headline finding that looks statistically robust until someone checks the underlying rows.
Collapse: where paleoclimate proxies join the comparison
Societal “collapse” — the breakdown of centralized administration, monumental construction, and dense settlement, as in the Classic Maya Terminal Classic period around 800–950 CE — has become a genuinely three-cornered problem, and a fourth data stream, paleoclimate proxy records, has become essential alongside the original three. Isotopic analysis of gypsum deposited in Lake Chichancanab, Mexico, allowed researchers to reconstruct rainfall and relative humidity through the Terminal Classic period with a precision not available from earlier lake-sediment methods, and the study concluded that annual rainfall fell by an average of around 50 percent, and by as much as 70 percent during peak drought years, with the most arid interval of the last two thousand years in that region falling between roughly 800 and 1000 CE — squarely overlapping the Classic Maya political collapse [10]. That is a fact-level paleoclimate finding, established through direct isotopic measurement of a physical sediment archive, comparable in kind to excavation and ancient DNA in that it measures a physical trace rather than coding a historical interpretation.
Whether drought caused the collapse, however, is an analytical claim that requires combining the climate record with excavated settlement-abandonment sequences region by region — and here the regional picture complicates a simple drought-causation story. Archaeological survey work on agricultural adaptation across the Maya lowlands found substantial regional variation in how communities responded to the same drying trend, with some adapting successfully for generations before eventual decline [11], and more recent site-specific paleoclimate work at Itzan found no local drought signal during the very period when that site’s population was declining in step with other, genuinely drought-affected communities elsewhere in the region — indicating that whatever caused Itzan’s decline, it was not simply local rainfall. Read together, these findings do not support “drought caused the collapse” as a single sufficient explanation; they support a mix of regionally variable climate stress, agricultural strategy, and — plausibly — the kind of political and trade interdependency between polities that only a comparative framework can characterize, since a single site’s abandonment sequence cannot on its own show whether its trouble originated locally or arrived from a stressed neighbor.
Regional diversity: a shared blind spot, addressed differently
All three approaches share a documented bias toward regions with long research traditions, funding, and political stability enabling sustained fieldwork — the Near East, the Aegean, Mesoamerica, and China are disproportionately excavated, disproportionately sequenced, and disproportionately well coded in comparative databases relative to, for instance, much of sub-Saharan Africa, Southeast Asia, or inner Eurasia, not because complex societies were less common there but because research infrastructure and access have been less continuous. Excavation addresses this only site by site, as new fieldwork programs open in under-studied regions. Ancient DNA has, in some respects, moved faster on this front recently, because sequencing does not require the decades of accumulated site-specific scholarship that deep excavation traditions took to build in classic regions — a single well-preserved skeletal sample from an under-studied region can enter the genomic record relatively quickly once sampling permission and preservation conditions allow it. Comparative databases are the most structurally exposed to this bias, since their statistical conclusions are only as globally representative as their case list, and a “global” pattern drawn from a sample skewed toward well-documented regions risks mistaking a research-availability artifact for a genuine cross-cultural regularity — precisely the kind of problem that made the Big Gods coding dispute so consequential.
What each approach cannot see, side by side
Held up against each other rather than described in isolation, the trade-offs are consistent across every topic above. Excavation gives the tightest, most directly falsifiable chronology at a single place but cannot establish whether a pattern generalizes beyond that place without additional survey or comparative work. Ancient DNA gives direct, unambiguous evidence of population movement, kinship, and pathogen history that no artifact typology can substitute for, but it is structurally silent on institutions, law, belief, and political authority — a genome cannot testify to who ruled. Comparative databases give the only genuine cross-case statistical test of hypotheses about state formation, religion, and collapse, but every one of those tests inherits the reliability, and any error, of the underlying case codings performed by people reading and interpreting the other two kinds of evidence — which is exactly the mechanism behind the one clearly documented failure in this article, the Big Gods retraction.
None of this supports ranking the three methods. The genuinely productive recent findings — Bronze Age plague lineages reshaping assumptions about Late Neolithic population decline, Göbekli Tepe reversing the assumed order of monumentality and agriculture, the corrected Big Gods reanalysis holding up a version of its original claim after the coding problems were fixed — all came from forcing two or more of these approaches to answer to each other’s evidence, not from trusting any single one in isolation.
Where the open questions sit, honestly
Several disagreements in this field remain genuinely unresolved rather than merely under-studied, and it is worth naming them as disagreements rather than implying a consensus exists. Whether drought was necessary, sufficient, or merely a contributing stressor in the Classic Maya collapse is actively disputed within the paleoclimate and archaeological communities, given the Itzan finding of local population decline without a matching local drought signal. Whether the corrected Big Gods pattern reflects a genuine cross-cultural regularity in the relationship between social scale and moralizing religion, or a residual artifact of which regions and periods happen to be well coded in Seshat, remains a live methodological question among the researchers who conducted the reanalysis themselves. And how much of early Neolithic Europe’s ancestry shift reflects wholesale population replacement of hunter-gatherers by incoming farmers, versus a more gradual, regionally variable admixture process, continues to be refined as more ancient genomes are sequenced from previously under-sampled regions, with different studies from Iberia, central Europe, and Britain reporting different admixture proportions [5] [6]. None of these three disagreements will be settled by any one of the three methods working alone.
A forward-looking note, held to a horizon
Looking out roughly ten years, it is a reasonable, evidence-grounded expectation — not a certainty — that ancient-DNA sample sizes from currently under-sampled regions (inner Eurasia, sub-Saharan Africa, Southeast Asia) will grow enough to test whether patterns established in the Near East and Europe generalize globally, because sequencing costs continue to fall and sampling programs in these regions have already begun. The assumptions behind that expectation are that political access for fieldwork continues in the relevant regions and that skeletal preservation conditions there are adequate for DNA recovery, which is not guaranteed everywhere. An observable indicator to watch for is the regional balance of new ancient genomes published in the major population-genetics journals over the next five years. The forecast would be disconfirmed if, a decade from now, the ancient-DNA record remains as geographically skewed toward the Near East and Europe as it is today, which would indicate that sequencing cost was never the binding constraint and that fieldwork access or preservation conditions are the deeper limiting factor instead.
Sources
This article draws on peer-reviewed genomics papers, the Seshat databank’s own published methodology, the Nature retraction and its follow-up reanalysis, and paleoclimate studies published in Science and PNAS, all verified live as of the date on this article. No statistic, quotation, or finding above is invented; each is attributed to the specific study that produced it.