A fossil does not arrive with a date attached, and a stretch of ancient DNA does not arrive labeled “3% Neanderthal.” Both are manufactured facts — the output of specific instruments, specific chemistry, and specific statistics, each of which can fail in specific ways. This briefing walks through three of the mechanisms that actually produce the numbers you read in a headline about human origins: how a fossil gets a date, how a genome gets pulled out of a bone that has been in the ground for tens of thousands of years, and how a proportion like “1–2% Neanderthal ancestry” gets estimated rather than measured directly.
Dating a fossil without dating the fossil
Fossilized bone itself usually cannot be radiometrically dated — the minerals that carry a datable radioactive clock in living tissue are mostly gone or altered by the time the bone has fossilized. So paleoanthropologists date the fossil’s context instead. The workhorse method for sites with a volcanic layer is argon-argon (⁴⁰Ar/³⁹Ar) dating, a refinement of potassium-argon dating [6]. Potassium-40 decays to argon-40 at a known, fixed rate; a volcanic mineral crystal — often sanidine feldspar — traps that argon-40 as it forms, starting a clock at the moment of eruption.
The argon-argon refinement irradiates the sample with neutrons first, converting a known fraction of potassium-39 into argon-39. That lets a mass spectrometer measure two argon isotopes in the same sample rather than measuring potassium and argon separately in different sub-samples, which removes a major source of error from older potassium-argon work [6]. In practice, a single sanidine grain is fused with a laser inside a vacuum chamber, releasing its argon, which is then measured directly against a known calibration standard such as Fish Canyon sanidine. The result is a date for the ash layer above or below the fossil, not the fossil itself — the fossil’s age is bracketed between two dated layers, a fact that is a source of real uncertainty whenever a site’s stratigraphy is disturbed or the fossil-bearing layer sits between widely spaced ash beds.
Not every site has volcanic ash. The 2013–2014 excavation of Homo naledi remains in South Africa’s Rising Star cave system had no volcanic layer to date, so the team instead combined uranium-thorium disequilibrium dating, electron spin resonance, optically stimulated luminescence, and paleomagnetic analysis on the flowstone, sediment, and teeth themselves, converging on an age of roughly 236,000–335,000 years for fossils with a strikingly primitive skeletal anatomy [5]. That fact — a small-brained, primitive-postcrania hominin persisting into a period when anatomically modern humans were already evolving elsewhere in Africa — is a documented finding, not an inference stacked on inference; it came from four independent dating methods converging on overlapping ranges, which is itself the discipline’s standard for confidence when no single method is decisive.
Pulling a genome out of a 50,000-year-old bone
Ancient DNA work starts from a problem: a bone that has spent tens of thousands of years in soil contains vanishingly little of its original DNA, that DNA is broken into short fragments and chemically damaged, and the sample is saturated with modern bacterial and human DNA that will swamp the signal if it gets in. The published extraction protocol used across major ancient-DNA labs addresses the first two problems with a silica-based purification method: bone or tooth powder is digested in an EDTA and proteinase-K buffer, and the released DNA — including fragments as short as 35 base pairs — is bound to silica in a chaotropic salt buffer optimized specifically to retain those short fragments, since standard silica columns designed for modern DNA lose most of what ancient DNA actually consists of [2].
Contamination is handled procedurally rather than chemically: work happens in a physically separate, positive-pressure clean room that no one who has worked with modern human DNA that day may enter, all reagents are UV-irradiated, and every batch of samples is processed alongside blank extraction controls that get carried through library preparation and sequencing exactly like real samples — if a control comes back with DNA in it, the batch’s results are suspect [2]. This is why the original Neandertal genome project reported its results only after building this kind of contamination-tracking pipeline: three individuals’ bone samples were extracted, converted into sequencing libraries, and in some cases enriched by hybridization capture for the mitochondrial genome before whole-genome shotgun sequencing was attempted [1]. The distinction between fact and inference matters here: that a set of DNA fragments were recovered and sequenced is a fact, verifiable against the raw sequencing reads; that this data represents authentic ancient Neandertal DNA rather than contamination is a claim supported by specific evidence — patterns of chemical damage (cytosine deamination) characteristic of degraded ancient DNA, fragment length distributions, and comparison against contemporaneous human reference panels — not an assumption.
Turning sequence overlap into a percentage
Once a Neandertal genome and a set of present-day human genomes exist side by side, “Neandertal ancestry” is not read off directly — it is estimated from patterns of allele sharing using formal population-genetic statistics. The foundational tool is the D-statistic, introduced alongside the draft Neandertal genome itself: for four populations arranged in a known tree, the method counts sites where allele patterns are inconsistent with simple vertical inheritance and asks whether that mismatch is symmetric (consistent with no gene flow) or skewed toward one population (consistent with admixture) [1] [3]. Applied to Neandertals, non-African, and African present-day genomes, the test found modern non-Africans significantly more similar to the Neandertal genome than Africans were — the original statistical basis for inferring interbreeding rather than shared ancient African variation [1].
A D-statistic tells you admixture happened; it does not by itself hand you a percentage. That comes from a related quantity, the f4-ratio, which compares one allele-sharing statistic against another calibrated one to produce a proportion — for example, dividing a statistic computed between a test population and Neandertal references by the same statistic computed against a second, more deeply diverged Neandertal or Denisovan reference, yielding an estimate such as the widely cited 1–2% figure for Neandertal ancestry in present-day non-Africans [3] [4]. A separate, computationally heavier approach — used by Sankararaman and colleagues to map where in the genome Neandertal ancestry sits rather than just how much exists — trains a probabilistic model on the patterns of variation expected under admixture and applies it locally across the genome of each individual, then aggregates the calls to find, for instance, that Neandertal ancestry is markedly depleted on the X chromosome and around genes highly expressed in testis, a pattern consistent with reduced fertility in early human-Neandertal hybrids [4]. That last inference — reduced hybrid fertility — is an analytical interpretation of an observed statistical pattern, not a directly observed fact; it is the best-supported explanation researchers have offered, and it is presented in the literature with that qualification.
Where the uncertainty actually lives
None of these three methods is a black box that spits out a settled truth. Argon-argon dates are only as good as the assumption that the dated layer was not later disturbed and genuinely brackets the fossil. Ancient DNA results are only as good as the contamination controls run alongside them, and for most of human prehistory, preservation conditions destroy DNA before it can be recovered at all — hot, humid environments in most of Africa and Southeast Asia have yielded very little ancient DNA compared to cold, dry, or cave contexts in higher latitudes, which biases the genomic record toward certain regions and populations. Admixture proportions depend on which reference populations and which statistical model are chosen, and estimates have shifted over successive papers as reference panels improved. A prediction sometimes voiced in press coverage — that better ancient-DNA recovery from tropical Africa will substantially revise current admixture estimates — is a testable forecast with a ten-year horizon at current sequencing-cost trajectories, whose disconfirmation condition would be several well-preserved tropical genomes yielding results consistent with existing estimates rather than overturning them. It is a plausible scenario, not a settled fact, and treating it as anything more would overstate what today’s methods can support.