Twelve flasks, one ancestor, and a rule with no exceptions
On February 24, 1988, Richard Lenski founded twelve populations of Escherichia coli from two ancestral clones of E. coli B that differ by a single neutral mutation: one clone can grow on arabinose (Ara+) and the other cannot (Ara−), a marker with no fitness effect in the experiment’s own environment but a visible one on the right plate, where Ara− cells form red colonies and Ara+ cells form white ones on tetrazolium-arabinose indicator agar [3]. Six populations were founded from each clone, so the marker exists purely to let two populations be told apart under a microscope’s worth of scrutiny, not because arabinose matters to anything that follows.
Every day since, a member of the lab has taken 1 percent of each population’s culture and diluted it into a flask of fresh medium — a hundredfold dilution that yields close to 6.64 generations of regrowth before the next transfer [3, 8]. The medium is DM25, a glucose-limited minimal medium holding 25 milligrams of glucose per liter as the growth-limiting resource, alongside roughly 1,700 micromolar citrate — about 280 milligrams per liter, on the order of eleven times the glucose by mass [3]. Flasks sit in a shaking incubator held at 37°C [14]. Every 500 generations, or about 75 days, a sample of each population is frozen in glycerol at −80°C, a step Lenski calls the experiment’s “frozen fossil record” [8].
The citrate deserves a full sentence rather than a parenthesis, because it is doing real work in this design and it becomes the hinge of the whole experiment thirty thousand generations later. E. coli has a complete tricarboxylic acid cycle and can metabolize citrate internally once it is inside the cell; what it cannot do, under the oxygen-rich conditions of a shaken flask, is get citrate across its own membrane, because the citrate transporter gene citT is not expressed aerobically — “the only known barrier to aerobic growth on citrate is its inability to transport citrate under oxic conditions,” as the 2008 paper documenting what eventually broke that barrier puts it [3]. Citrate is in DM25 as a chelating agent, keeping iron in a soluble, bioavailable form that the cells can take up through their own separate ferric-citrate transport machinery — present, in other words, as a mineral-handling additive, not as food. For the first three decades of the experiment this distinction held perfectly: a resource sat in every flask, at roughly eleven times the concentration of the sugar the bacteria were actually competing over, and nothing used it.
The marker itself is worth pausing on, because it is the clearest sign that the design was built to answer questions its founder had not yet asked. Arabinose use has no measurable effect on fitness inside DM25; it exists solely so that two populations, or an ancestor and its own descendant, can be told apart on sight after they have been mixed in the same flask and grown side by side for a day. Every competitive fitness number in this literature — including the ones anchoring the citrate story below — depends on that visible, cost-free label holding steady across tens of thousands of generations. Twelve populations is itself a deliberate choice rather than a round number: enough independent lineages to ask whether a given outcome is a rule or an accident, too many for any one experimenter to have hand-picked a favorable result after the fact. At roughly 6.64 generations per day, the arithmetic of the design means that what a large-bodied, slowly reproducing organism would need on the order of a million years to accumulate in generation count, a shaken flask of E. coli can pass through in about thirty-three years — the gap that makes the LTEE legible within a single research career at all.
Lenski’s first published report on the experiment, at 2,000 generations, already contained the pattern that would define everything after it: mean fitness rose across all twelve populations, and the rate of that rise slowed as the generations accumulated, with significant divergence between populations but no detectable heritable variation in fitness within any one of them [1]. Fox and Lenski’s own retrospective on the design explains the choice behind all of this as deliberate: “I had made a strategic decision to make the environment of the LTEE very simple to eliminate, or at least reduce, certain complications” [8]. One strain, one flask type, one sugar, one transfer ritual, repeated without exception for what is now approaching four decades. The austerity is the instrument.
The freezer, not the notebook, is the record
What makes the LTEE more than a very long anecdote is that its history is not only written down — it is stored, viable, and re-runnable. A vial from generation 2,000 is not a data point about the past; it is a living culture that can be thawed today, revived, and run head-to-head against a competitor from generation 50,000, or against the original ancestor, in the same flask, under the same conditions, this afternoon. Every fitness estimate in this literature depends on that possibility. A competitive fitness assay works by reviving a marked competitor — commonly the frozen ancestor itself — and a test population, mixing them in one flask of DM25 for a day’s growth, then plating the mixture on tetrazolium-arabinose agar and counting the red and white colonies to estimate each strain’s relative growth rate, using the neutral Ara marker described above purely as a way to tell the two strains apart under the count [3].
The frozen archive is not a hypothetical convenience; it has been used to restart the actual, ongoing experiment at least once. When the SARS-CoV-2 pandemic closed Lenski’s Michigan State laboratory in March 2020, all twelve LTEE lines were frozen in place. On September 22, 2020, lab member Devin Lake restarted the experiment from those same frozen samples, at generation 73,000, putting the populations “back in their home-sweet-homes: Erlenmeyer flasks with DM25 medium” in a shaking incubator at 37°C [14]. The experiment did not resume from a fresh culture or a best approximation; it resumed from the exact frozen state it had been in six months earlier, because that state had been kept viable for precisely this kind of interruption. The fossil-record metaphor is, in this one instance, literally what happened: the living experiment was rebooted from its own preserved past.
Against that record, Wiser, Ribeck and Lenski fit fitness trajectories from all twelve populations across 50,000 generations against two competing functional shapes for how relative fitness might grow with time. One is a form that saturates — fitness rising quickly at first and then flattening toward a fixed asymptote, the kind of ceiling a “diminishing returns” reading of Lenski’s 1991 result would predict. The other is a power-law form, with no fixed limit: growth continues indefinitely, at an ever-slowing rate that nonetheless never reaches zero. Schematically, the contrast is between something like
which approaches a fixed value of
for an exponent
The trajectory has kept behaving that way. By 2024, the LTEE populations were reported to have reached “all-time peak fitness,” having accumulated much of their earliest improvement in three discrete steps of roughly 10 percent each, and as of a report published in December 2025 there were “no plans to draw the experiment to a close” [13]. Thirty-eight years in, the flasks are still gaining — more slowly than in year one, and still gaining.
One population learned to eat what was already in the flask
For more than 30,000 generations, no population evolved the ability to use the citrate sitting in every flask, despite each one testing on the order of billions of mutations across that span [3]. Then, in the population designated Ara-3, between generations 31,000 and 31,500, a variant appeared that could grow aerobically on citrate — a citrate-positive, or Cit+, phenotype — and its emergence visibly changed the population, expanding both its total size and its diversity as the new resource opened up additional carrying capacity [3].
Blount, Borland and Lenski then asked a harder question than “did it evolve”: could it have evolved this early, in principle, or did something about the population’s own history have to happen first? They revived frozen clones from many points across Ara-3’s history and let them evolve again, in triplicate or more, to see whether Cit+ would reappear and from which starting points. In the first of three replay experiments, 72 replay populations — six independent replicates founded from each of several historical generations — were evolved for roughly 3,700 generations apiece; in a second, 340 separate cultures were plated directly onto citrate medium and incubated for 59 days; in a third, on the order of
Blount, Barrick, Davidson and Lenski’s follow-up genome sequencing filled in the mechanism and the population structure underneath it. Three distinct clades, labeled C1, C2 and C3, had already diverged and coexisted within Ara-3 well before Cit+ appeared: C1 split from the common ancestor of C2 and C3 before generation 15,000, and C2 and C3 had themselves diverged by around generation 20,000 [4]. Cit+ originated within one of these clades through a tandem duplication of a roughly 2,933-base-pair segment containing the citT gene and a neighboring gene, rna — a duplication that placed an aerobically active promoter next to citT, driving expression of a transporter that had been present all along but silent under oxygen [4]. The paper’s own description of the general phenomenon — “promoter capture and altered gene regulation… mediating the exaptation events that often underlie evolutionary innovations” — names what happened precisely: no new gene was invented, an existing one was switched on in a context where it had never been read before.
The new module, an rnk-citT duplication, then refined itself further: from an initial two-copy tandem array around generation 31,500, copy number climbed to three-, four- and even nine-copy arrays by generations 32,000 to 33,000, before settling toward a four-copy configuration that gave a robust Cit+ phenotype and let citrate use expand to dominate the population’s biomass by roughly generation 33,000 [4]. When the authors resequenced 19 independently re-evolved Cit+ mutants recovered from later replay experiments, they found the same regulatory trick — a promoter placed in front of citT — achieved through several distinct mutational routes: duplications with different boundaries, insertion-sequence-mediated rearrangements, inversions, deletions [4]. The specific genetic route was not fixed, but the functional solution, promoter capture at citT, recurred every time. A separate, smaller ecological consequence followed: the original Cit− cells did not vanish once Cit+ took over the citrate niche, but persisted alongside it in a cross-feeding relationship, living off the C4-dicarboxylate compounds that Cit+ cells excrete while growing on citrate [3] — a single-resource flask that had, without anyone designing it to, generated two coexisting ecological types.
This account has a serious challenger, and the register of the disagreement is worth stating precisely rather than picking a side by omission. Van Hofwegen, Hovde and Minnich set out to test whether the LTEE’s numbers reflected an intrinsically rare, historically contingent event or an artifact of the LTEE’s own protocol. Working outside the LTEE, using the ancestral strain but direct selection on solid citrate medium rather than the LTEE’s daily hundredfold liquid dilution, they isolated 46 independent Cit+ mutants, with potentiation and initial actualization occurring in as few as 12 generations and phenotypic refinement complete within about 100 generations — roughly three orders of magnitude faster than the 33,000 generations the trait took inside the LTEE [7]. Their proposed mechanism again centered on citT, amplified about fourfold, together with a roughly twofold amplification of a second transporter gene, dctA, needed to recapture succinate; their own summary is blunt that “no new genetic information (novel gene function)” evolved, only expanded expression of transporters already present [7]. Their argument is that the LTEE’s specific numbers — the 30,000-generation wait, the four-in-dozens replay success rate — are largely artifacts of that protocol: daily hundredfold dilution flushes out low-frequency intermediate mutants before they can be enriched, glucose in the medium represses citT transcription through catabolite repression, and the ancestral strain used in the LTEE carries a defective dcuS that hampers succinate sensing, all of which work against Cit+ appearing quickly under LTEE conditions specifically [7].
Both results can be true at once, and the honest statement is not that one side is simply wrong. The replay experiment’s core finding — that only genetic backgrounds from late in Ara-3’s own specific history ever produced Cit+ under the LTEE’s own fixed protocol, however many cells were screened from earlier ones — stands regardless of what a different selection scheme can achieve [3]. What Van Hofwegen and colleagues show is that changing the protocol changes the tempo dramatically, which means the specific numbers attached to “how contingent” and “how rare” are protocol-dependent rather than fixed properties of the mutation itself [7]. What the dispute does not touch is the direction of the underlying biology: both papers agree that no fundamentally new gene function was created, and both center the same two genes, citT and dctA, as the levers involved. The disagreement is about how much of the thirty-thousand-generation wait was biology and how much was one specific experiment’s own dilution regime — a question about magnitude and cause, not about whether the innovation happened.
The genome agrees with the flask, gene by gene
Fitness assays and replay experiments describe phenotypes; full-genome sequencing describes what changed underneath them, and by the time Tenaillon and colleagues sequenced 264 complete genomes spanning all twelve populations through 50,000 generations, the two pictures matched in a way that is hard to attribute to chance. Across those genomes they tallied 14,572 point mutations, roughly 500 insertion-sequence-element insertions, 726 small deletions and 1,132 small insertions of 50 base pairs or fewer, and 267 larger deletions alongside 45 duplications above that size [5]. Fifty-seven genes were hit by two or more independent mutations across the twelve lineages — lineages that have shared no genetic contact since 1988 — and those 57 genes alone carried 50.1 percent of all nonsynonymous mutations recovered, despite making up only 2.1 percent of the coding genome, a concentration the authors report as far beyond chance (Z = 25.5, p < 10⁻¹⁴³) [5]. Twelve independent experiments, run from the same starting clone under the same daily rule, kept finding the same targets.
Not every lineage evolved at the same rate, and the reason is itself now well characterized. Six of the twelve populations evolved elevated mutation rates — so-called mutator, or hypermutator, phenotypes, typically through defects in DNA repair — and by 50,000 generations those six populations alone accounted for 96.5 percent of all point mutations recovered across the whole experiment, at a rate on the order of a hundredfold above the ancestral mutation rate [5]. A higher mutation supply did not translate into a proportionally higher fitness ceiling; instead, the paper’s model shows the fraction of new mutations that are beneficial declining as fitness rises, with neutral mutations accumulating at a roughly constant background rate regardless of a lineage’s mutator status [5]. More mutations bought more genetic noise and a faster accumulation of hitchhiking neutral variants, not a different long-run destination.
Good, McDonald, Barrick, Lenski and Desai extended the picture to 60,000 generations using metagenomic sequencing of whole-population samples taken every 500 generations, rather than sequencing single clones, which let them watch competing lineages within a population rather than assuming one eventually swept to fixation. In nine of the twelve populations, multiple distinct genetic clades coexisted for more than 10,000 generations, in many cases persisting all the way to generation 60,000 [6]. Selection did not settle on a fixed list of targets and then stop: the gene hslU was mutated frequently early in the experiment and almost never late, while atoS showed close to the opposite pattern, mutated rarely early and repeatedly late [6]. The paper counts roughly sixteen “missed opportunities” — cases where one population never acquired a mutation in a gene that had been hit repeatedly and independently in its eleven sister populations, evidence that which mutations are available and useful to a lineage depends on the particular genetic background that lineage has already accumulated, not on a fixed target list common to all twelve [6]. Parallelism at the level of genes and contingency at the level of alleles and timing are not competing explanations here; they are two grains of the same result. The same functional targets keep recurring across independent populations, and which specific mutation gets used, in what order, and whether a given population reaches it at all, depends on that population’s own prior history — precisely the logic the Cit+ replay experiment demonstrated at the scale of one trait, now visible genome-wide.
A 2026 reanalysis of gene-level evolutionary rates across the LTEE lineages, tracking synonymous and nonsynonymous substitution patterns gene by gene, adds a temporal axis to the same parallelism: growth-related genes tended to evolve early in the experiment, while survival-related genes evolved later, and named recurrent targets across populations include rbs, spoT, topA, fis, malT and pykF [15]. The same analysis confirms the split the earlier genome papers had already found at the level of raw mutation counts: mutator populations carried consistently higher numbers of genes bearing both nonsynonymous and synonymous substitutions than the six nonmutator populations did, even as both groups showed the same general decline in gene-level evolutionary rate over time [15]. A higher mutation supply changed how many genes accumulated detectable change, not the direction the change was heading in — the same conclusion Tenaillon and colleagues had reached from the mutation-counting side ten years earlier, now visible from the gene-rate side as well.
The experiment has now outlived two of its own labs
By the summer of 2022, the twelve LTEE populations stood at 75,000 generations, and Lenski — then 65 and preparing to close his Michigan State laboratory — handed daily responsibility for the experiment to Jeffrey Barrick, a former postdoctoral researcher on the LTEE from 2006 to 2010, then running his own lab at the University of Texas at Austin [11, 12, 10]. The transfer was as literal as the science: frozen vials representing 34 years of accumulated history were packed and shipped from Michigan to Texas, and Lenski sent the populations off publicly with a line that captures the whole design in one sentence — “Bon voyage, #LTEE! Enjoy your new locale, even if your Erlenmeyer flask homes and DM25 diets are exactly the same as you’ve been evolving in” for 75,000 generations, encouraging the bacteria to “keep on evolving” [12].
At UT Austin, Barrick restarted the daily transfers and described the physical footprint of what had arrived: the entire primary archive of 500-generation-interval frozen samples, spanning 1988 to 2022, still fit into “about half of a standard table-sized chest freezer” [11]. His stated reason for taking on the commitment was about time itself as the experimental variable: “Time is really important for seeing evolution in action. The longer the experiment, the more interesting things you can see develop” [11]. Three years later, the institutional story added a second turn that the original 2022 coverage could not have anticipated: in 2025, Barrick was hired by Michigan State University’s Department of Microbiology, Genetics, and Immunology as a Hannah Distinguished Professor, and that August the LTEE — freezer archive, active flasks and all — moved back to Michigan State, the institution where it began [16]. The experiment has now changed its formal institutional address twice in three years without a single day of its daily 1:100 transfer protocol changing at all.
By the official project record, the twelve populations passed 80,000 generations during Barrick’s tenure, a milestone reported as reached in 2024 [9, 13]. As of the most recent public reporting checked for this article, in December 2025, the project had no announced plan to end the experiment [13]. What has proven durable across two campuses, two principal investigators, and one global pandemic is not the address or even the lab’s personnel — it is the protocol itself: the same ancestral lineage, the same DM25 recipe, the same hundredfold daily dilution, the same 500-generation freeze. The design’s austerity, built in 1988 to remove confounds from the biology, turned out incidentally to be the thing best suited to survive an institution’s own turnover.
What changed, what didn’t, and what this design can’t tell you
It is true, and worth saying plainly rather than around, that the twelve LTEE populations are still Escherichia coli. No population has crossed into a different genus, grown a new organelle, or become multicellular; the objection that “bacteria are still bacteria” is not wrong as a description, only wrong as a rebuttal, because nothing about the experiment’s evidentiary claims depends on that not being true. State precisely what the record actually supports. Measured competitive fitness has risen for 80,000-plus generations with no observed ceiling, tracking a power-law model against a saturating alternative that the data reject [2]. One population evolved a categorically new resource-use phenotype — aerobic growth on a compound present in the medium from day one — through a specific, sequenced mechanism of promoter capture, arrived at independently by at least nineteen separately observed mutational routes converging on the same regulatory logic [4]. Genome sequencing across all twelve lineages finds the same 57 genes recurrently and independently mutated far beyond chance, alongside population-specific hypermutator phenotypes, coexisting clades sustained for tens of thousands of generations, and a documented shift in which genes are usable targets as a lineage’s own background changes [5, 6]. None of that requires, or produces, a new species in any sense a zoologist would recognize, and none of it needs to, to count as open-ended adaptation, measured contingency, and innovation by co-option rather than de novo invention.
The design’s limits are just as specific as its results, and they follow directly from the same choices that made the results measurable. The LTEE is asexual: twelve clonal lineages accumulate mutations and occasionally lose them to selection or drift, but there is no recombination bringing together beneficial variants that arose in different populations, which is precisely why clonal interference — competing beneficial lineages within one population, unable to combine — is such a persistent feature of the molecular record here [6]. It runs in one environment: a single limiting sugar at a fixed low concentration, a fixed temperature, a fixed daily disturbance regime, which is exactly the constraint that let Lenski isolate a clean rate-of-adaptation signal in 1991 and exactly the reason the experiment says little by itself about adaptation to fluctuating, multi-resource, or biotically complex environments [1]. And it was designed, deliberately, without ecology: no predators, no competitors from outside the flask, no spatial structure. What the Cit+ story shows is that this constraint is not airtight even inside the experiment’s own rules — a single resource, unintentionally, generated a second, cross-feeding ecotype within one population [3] — but an accidental two-member interaction inside one flask is not evidence about community ecology, and treating it as such would overstate what one emergent cross-feeding pair can support. The clonal interference that Good and colleagues documented in nine of twelve populations cuts the same way: multiple lineages competing inside one population for over 10,000 generations is a genuine population-genetic phenomenon with real consequences for which mutations fix, but it is competition between genotypes of one species over one resource, not the trophic structure, spatial patchiness or multispecies interdependence that “ecology” usually names in a wild system [6]. The LTEE was built to isolate adaptation from ecology as cleanly as a wet-lab experiment can, and it mostly succeeds; where it doesn’t — Ara-3’s cross-feeding pair chief among the exceptions — the exception is small enough to prove the general containment rather than undermine it.
A modest, falsifiable claim follows from the power-law result, and it is worth stating with its own horizon and its own way of being wrong. On the assumption that the current protocol and environment continue unchanged and that no population undergoes a further Cit±scale innovation that resets its competitive baseline, the power-law fit from Wiser and colleagues predicts that mean fitness gains per additional 10,000 generations should continue shrinking but remain measurably positive through at least the LTEE’s second century of generations — a horizon on the order of 100,000 to 150,000 generations from the 1988 start, which the current 6.64-generations-per-day pace would reach within the next two to three decades [2]. That prediction would be disconfirmed by a sustained, multi-thousand-generation plateau in competitive fitness against the ancestor in any population not attributable to a measurement artifact — precisely the saturating alternative the 2013 paper’s own model comparison already rejected once, using data through 50,000 generations [2]. Whether it holds through 150,000 is, at time of writing, still an open question the freezer is positioned to answer whenever someone asks it.