A moth that turned black taught a district how fast selection can move

In 1848, a collector named R. S. Edleston caught a black specimen of the peppered moth, Biston betularia, near Manchester — a form so unlike the common pale, speckled type that it was given its own name, carbonaria. He did not publish the capture until 1864, sixteen years later, by which time it no longer looked like a curiosity. It looked like the start of a trend [9]. By 1895, on the figure Bruce Grant and Lawrence Wiseman report from the historical record, about ninety-eight per cent of the peppered moths collected near Manchester were the black form [9]. A single aberrant moth had gone from unique to a near-total replacement of the wild type across an entire industrial district in roughly two generations of a species that in Britain has one generation a year, meaning that the whole reversal fits inside a human lifetime.

The species itself gives a name to the phenomenon it demonstrates: melanism driven by industry. Manchester in the middle of the nineteenth century was burning coal at a rate that blackened stone, killed lichen and coated tree bark in soot, and the region’s peppered moths spend the day resting motionless on bark, relying on camouflage against exactly that surface to avoid being eaten by birds that hunt visually. A moth that is pale and mottled is nearly invisible against lichen-covered, unpolluted bark and glaringly obvious against soot; a black moth is the reverse. If birds really do find and eat moths that stand out, and if pollution really had changed which morph stood out, the frequency of carbonaria should have tracked the pollution, district by district, on a timescale of decades rather than centuries. Laurence Cook’s 2003 review in the Quarterly Review of Biology describes exactly that geography taking shape by the middle of the twentieth century: a steep cline running from northeastern England, where melanic frequency reached ninety per cent or higher, down to five per cent or less in unpolluted North Wales [5]. The pattern was not confined to one city or one country. Grant, Owen and Clarke documented an almost identical rise in industrial Michigan, reaching above ninety per cent by 1959, tracking local sulphur dioxide and particulate pollution on the same logic as the British case [10]. Two populations of the same species, on two continents, converged on the same visible trait for what looked like the same visible reason. That convergence is what made Biston betularia worth turning into an experiment rather than leaving it as a natural-history anecdote.

An observation of correlated change is not itself evidence of predation, and the crucial gap in the nineteenth-century record was direct: nobody had actually watched a bird eat a moth off a trunk and shown that which moth got eaten depended on which trunk it was sitting on. Closing that gap, deliberately and with a falsifiable design, is what the next thirty years of the case are about — and so, eventually, is reopening it.

ADVERTISEMENT

The cline itself was never a clean step function between a fully melanic industrial zone and a fully pale rural one, and the reason matters for how the case was tested later. Cook and Saccheri’s retrospective credits a 1972 mark-recapture study by Bishop with the first real estimate of how far individual moths disperse in a generation, putting average movement for male moths at roughly two kilometres a night — enough gene flow, sustained over decades, to blur what a purely local selection pressure would otherwise render as a sharp border and instead produce the smooth geographic gradient Cook’s later cline data actually show [2]. A trait under strong local selection but subject to ordinary dispersal should look exactly like this: high in the core of the polluted zone, low far outside it, and a gradual slope in between where immigration from one side keeps depositing the “wrong” morph into the other. That is a second, independent prediction the migration-selection framework makes, distinct from the trunk-by-trunk predation claim, and it is one the geography had already confirmed before Kettlewell ever set a trap.

A single antique pinned peppered moth specimen held up to raking light on a cork board, its small aged locality label caught mid-turn as it is checked against the drawer's other entries
Figure 1. The first captured black moth was caught near Manchester in 1848 and not reported for sixteen years; by 1895 near-fixation had already overtaken the district it was found in.Image prompt and art direction by Brecht Corbeel; generation pending.

Kettlewell’s traps measured a real, direction-dependent difference in survival

Bernard Kettlewell, working from Oxford in the 1950s, designed the experiment the correlation demanded: mark moths of both morphs, release them into real woodland, and see which morph a trap recovered more of after the birds had had their chance. The logic of mark-release-recapture is that recapture rate is a proxy for survival between release and recapture; if predation is morph-blind, both morphs should come back at similar rates regardless of habitat, and if predation tracks camouflage, the rates should flip between polluted and unpolluted sites.

He ran the design in two contrasting places. In sooted woodland near Birmingham, in an industrial district where the bark itself had been darkened by decades of smoke, he released marked moths of both morphs and recovered carbonaria at roughly twice the rate at which he recovered typica. In unpolluted woodland in Dorset, where the bark still carried its lichen, he ran the same design and recovered typica at roughly twice the rate at which he recovered carbonaria — the ratio inverted, in the direction the camouflage hypothesis specifically predicted rather than in some direction consistent with any story at all [1, 2]. Laurence Cook and Ilik Saccheri’s 2012 retrospective describes Kettlewell’s own experiments as delivering the first convincing demonstration that birds eat the moths and can do so selectively, and characterizes the melanic advantage in industrial sites at high carbonaria frequency as running as high as two to one over the typical form [2]. That two-to-one figure recurs on both sides of Kettlewell’s design, which is exactly the symmetry a real, camouflage-driven predation effect should produce and an artefact of his methods should not.

What the experiment did not do is prove that predation was the whole explanation for the geographic cline, or that Kettlewell’s particular way of running it was beyond methodological reproach. Both of those were fair questions in 1955, and both remained fair questions for the next forty years, because a mark-recapture experiment estimates a difference in survival, not a mechanism, and it estimates that difference under whatever conditions the experimenter actually created — release density, time of day, the substrate the moths were set on before they dispersed. Kettlewell’s contemporaries accepted the headline result quickly, and did so, in part, because it had a coherent alternative geography (Dorset versus Birmingham) built into its own design; what took another half-century to settle was whether the specific numbers he reported reflected the wild population’s real experience or the experiment’s own artificial setup.

It also mattered that Kettlewell did not run the comparison once. Cook and Saccheri’s account traces a sequence of his own follow-up trials through the 1950s, and then two further coordinated national surveys, in 1958 and in 1965, that other researchers used as a fixed baseline against which to measure how later work compared [2]. Repetition across years and sites is exactly what a genuine effect should tolerate and an experimental artefact should not: a fluke of one release date, one weather pattern, or one predator’s temporary preference should wash out under replication, while a real camouflage-driven difference should keep reappearing in the same direction wherever the same contrast in bark colour exists. It kept reappearing. The open question after Kettlewell was never really “did birds eat more of the conspicuous morph in these trials” — the recapture asymmetry was reproduced too often for that to be in serious doubt — but whether the conditions of the trials themselves had been close enough to the moths’ actual wild behaviour for the measured size of the effect to be trusted at face value.

ADVERTISEMENT
A portable mercury-vapour moth trap on a woodland-edge table at dawn, its perspex baffle lifted and one egg-box liner half-slid out with pale and dark peppered moths still settled among the cells
Figure 2. Kettlewell's mark-release-recapture design measured which morph came back to the trap: in the sooted wood near Birmingham the dark form returned at roughly twice the rate of the pale one, and in unpolluted Dorset woodland the ratio reversed.Image prompt and art direction by Brecht Corbeel; generation pending.

Majerus’s 1998 critique targeted the setup, not the mechanism

Michael Majerus, a Cambridge geneticist who had spent decades working on the same species, published Melanism: Evolution in Action in 1998, a book that devoted a substantial portion of its length to reassessing exactly this case on its twenty-fifth anniversary. His central objection was about ecological realism rather than about whether birds eat moths at all: peppered moths in the wild rarely rest in the open on vertical trunk faces, the position Kettlewell and most subsequent experimenters had used to present moths to predators, and Majerus argued that releasing moths at artificially high density, at an unnatural time of day, onto an unnatural resting surface, could inflate or distort a predation signal in ways that had nothing to do with what happens to the wild population — a set of concerns later summarized as the “bird table effect” [11]. This is a methodological critique in the ordinary sense that field ecology runs on: it does not dispute that birds are visual predators, that camouflage affects detectability, or that the geographic cline exists. It disputes whether one specific experimental design measured the wild process cleanly enough to support the numbers attached to it.

The critique escaped its own scope almost immediately, through two separate and importantly different channels. The geneticist Jerry Coyne reviewed Majerus’s book in Nature and, reading the density and resting-site concerns as sufficiently serious, wrote that biologists should for the time being set Biston aside as a well-understood textbook example pending better data — a scientist’s provisional and falsifiable statement of doubt, aimed at the specific experimental record, not at natural selection as a phenomenon [11]. That sentence was then lifted out of its qualified context by anti-evolution advocacy literature and circulated as though a leading evolutionary biologist had conceded that the whole case for natural selection collapsed with it — a use Coyne did not intend and explicitly disowned [11]. The second channel did far more damage to the historical record: the science journalist Judith Hooper’s 2002 book Of Moths and Men went beyond Majerus’s methodological argument to insinuate that Kettlewell had committed outright fraud, a claim built substantially on secondhand and circumstantial material. The response from working geneticists who had known Kettlewell was blunt. Bryan Clarke, who had worked alongside him at Oxford, called the book a treasury of insinuation worthy of an unscrupulous newspaper, and a subsequent point-by-point academic assessment concluded that none of Hooper’s arguments withstood careful scrutiny [11].

It is worth being precise about what each of these three things was, because they are not interchangeable and folding them together is how the story gets told badly in both directions. Majerus’s critique was a scientist identifying a genuine, testable weakness in a forty-year-old experimental design. Coyne’s review was a provisional judgment, later misquoted, about how much weight that specific design could still bear. Hooper’s book was an accusation of scientific misconduct that the specialists who examined it did not accept. Creationist literature then used the second to imply the third, and used both to imply that the underlying biology — the geographic cline, the pollution correlation, the camouflage mechanism itself — had been discredited, which none of Majerus’s own published work ever claimed.

A fourth strand, distinct again, ran through popular textbook criticism rather than through the primary research: the claim, pressed by the creationist author Jonathan Wells among others, that classroom photographs of moths resting on trunks had been staged using dead or immobilized specimens deliberately posed for the camera, and that this staging amounted to teaching a fabrication rather than a finding. The National Center for Science Education, reviewing that specific claim, distinguishes it sharply from Majerus’s actual argument: Wells’ complaint was aimed at illustrative photography and at the sluggishness of moths released in daylight for some early photographs, not at the mark-recapture data itself, and subsequent controlled releases timed to the moths’ natural dawn activity reproduced the same differential recapture pattern regardless of how any textbook picture had been staged [1]. A misleading photograph and a flawed experiment are different failures with different remedies, and neither one is evidence against the other; the persistent public confusion of “the picture in the textbook was staged” with “the result in the paper is false” is itself a small case study in how a legitimate objection to illustration can be laundered into a much larger and unsupported claim about the underlying science.

Two matched reference panels of real tree bark on a collection-room bench, one furred with pale lichen and one darkened with soot, a pinned moth specimen being tried against a bark crevice rather than the open trunk face
Figure 3. Majerus's 1998 critique was that Kettlewell's moths were usually set on open trunk faces they rarely chose in the wild; the objection was about where the moths were placed, not about whether birds ate them.Image prompt and art direction by Brecht Corbeel; generation pending.

Majerus spent his last seven years answering his own critique

What makes this case unusual, and is the reason it belongs in a discussion of self-correction rather than merely of controversy, is what its most credible critic did next. Rather than leaving the methodological objection as a standing question mark over the textbooks, Majerus designed and ran an experiment built specifically to remove every artefact he himself had identified in Kettlewell’s protocol. Between 2001 and 2007, in his own garden at Springfield, near Coton in Cambridgeshire, he released a total of 4,864 moths, allowing them to settle at the natural density and natural morph frequency for that location rather than at an experimenter-chosen ratio, and recording the outcome as the moths themselves disposed of it rather than intervening in where they rested [3, 2]. He also went back to the resting-site question directly, surveying where wild moths actually settle rather than assuming an answer either way. The posthumous analysis of his data reports that thirty-five per cent of the naturally settled moths he recorded rested on tree trunks — the very site Majerus’s own 1998 book had flagged as artificially over-used by earlier experimenters — which meant the trunk-resting design was not the fabricated setup his critique might have implied, only an over-simplified one [3].

Majerus died in January 2009, before he had finished analysing all six years of releases himself [11]. Laurence Cook, Bruce Grant, Ilik Saccheri and James Mallet completed the work and published it in 2012 in Biology Letters, under a title that names whose experiment it was: “the last experiment of Michael Majerus.” Their analysis found strong, statistically supported evidence of overall selection against the melanic form at that site and in that decade, with a daily selection coefficient against carbonaria of approximately 0.091, a ninety-five per cent confidence interval of 0.028 to 0.157 — daily survival for the melanic form running at roughly ninety-one per cent of the typical form’s survival across the study [3]. The authors’ own conclusion is that a differential of that magnitude and direction is sufficient to explain the rapid decline of melanism that had by then been under way in Britain for decades [3]. The design also let the analysis rule predators in and out by type rather than assume them: the recorded pattern of disappearance was consistent with daytime, visually hunting birds and specifically inconsistent with bats, which hunt nocturnally by echolocation and would not be expected to discriminate wing colour at all, and the data showed no corresponding bat-driven signal, which is exactly the negative control a genuinely visual, avian predation mechanism should produce [11].

ADVERTISEMENT

This is not the same number, measured the same way, as Kettlewell’s original recapture ratios — it comes from a different decade, a different site, and natural rather than experimenter-chosen density and placement — and that is precisely what makes it a real test rather than a repeat performance. Majerus set out, on his own account, to find out whether his critique of Kettlewell would survive contact with a design built to answer it. It did not survive; differential bird predation, in the direction and rough magnitude the textbook case had always claimed, held up under exactly the conditions Majerus had argued were missing from the original.

An open ruled data ledger on a collection-room bench, columns for date, site and morph filled down the page, the current row only half filled in beside a hand-tally counter reading short of a round number
Figure 4. Between 2001 and 2007 Michael Majerus released 4,864 moths at natural density onto natural resting sites in his own garden near Cambridge, the design built to answer his own 1998 critique on its own terms.Image prompt and art direction by Brecht Corbeel; generation pending.

The mutation was dated to about 1819 and the camouflage was measured, not asserted

Behavioural ecology settled what was happening; population and molecular genetics, over the following decade, settled how and when. Arjen van 't Hof and Ilik Saccheri’s group first mapped the genetic basis of the carbonaria trait in 2011, publishing in Science a fine-scale association analysis that localized melanism to a single small genomic region and found only one haplotype associated with the melanic phenotype across the sampled population, carrying the signature of a single, recent, strongly selected mutational event rather than several independent mutations converging on the same trait [6]. That is a strong, specific and falsifiable claim in its own right: a trait as visually dramatic and as widespread as carbonaria could in principle have arisen more than once, in different lineages, at different times; the 2011 mapping data said it had not.

The logic behind that inference is worth stating plainly, because it is what lets a population-genetic study answer a historical question. If the black form had arisen independently many times — in different trees, different decades, different corners of Britain — each independent origin would in principle carry its own distinct stretch of surrounding DNA, tagged by whatever mutations happened to be sitting nearby on the chromosome where the change first occurred. Van 't Hof’s 2011 mapping instead found essentially one core haplotype shared across melanic moths sampled from across the country, the genomic signature of a single mutational event that then spread by ordinary inheritance and selection rather than of the trait re-evolving over and over [6]. The same locus turned out to sit in a region already known, from work on Heliconius butterflies, to control wing colour pattern in a completely different insect lineage — a coincidence of location that made the peppered moth’s melanism look, even before the causal mutation itself was found, like a case of natural selection reusing a genomic address that development had already made unusually easy to mutate into a colour switch [2].

Five years later, the same group identified the mutation itself. Van 't Hof and colleagues, publishing in Nature in 2016, showed that the melanism-causing event was the insertion of a large transposable element into the first intron of a gene called cortex, and that this insertion increases the abundance of a cortex transcript whose protein product functions in cell-cycle regulation during the early development of the wing disc [7]. Statistical inference from the pattern of recombination in wild carbonaria haplotypes — essentially, using how much the original insertion’s surrounding DNA had been shuffled by subsequent generations of recombination as a molecular clock — dated the transposition event to approximately 1819, a period, as the authors note, consistent with the early acceleration of British industrial coal burning and squarely inside the historical window in which the phenotype was first noticed and then swept toward fixation [7]. A subsequent genome assembly of the species, published as an open resource in 2022 and spanning some 405 megabases across a nearly complete set of chromosomes, confirms cortex’s role and situates it within a locus that recurs across other Lepidoptera as a hotspot for colour-pattern evolution, meaning the peppered moth’s mutation used a genomic mechanism natural selection appears to have reused many times over in this insect order, not a one-off fluke of one species [12]. That assembly also means the case is no longer methodologically frozen in the 1950s: the same reference genome that let van 't Hof’s group date one insertion to 1819 is now the standing resource against which any future question about this species — a new colour form, a new resistant population, a new local adaptation — can be checked directly against DNA rather than argued from morphology and geography alone, which is itself a kind of insurance against the case ever again resting on evidence as hard to re-examine as a fifty-year-old field notebook.

The remaining leg of the case was to quantify the camouflage itself rather than assert it by eye, because “birds can obviously see the difference” is exactly the kind of intuitive claim that a rigorous critique is entitled to ask for evidence on. Olivia Walton and Martin Stevens did that in 2018, in Communications Biology, using digital image analysis calibrated to blue tit colour vision to measure how well each morph’s wing pattern matched real bark and lichen backgrounds in the actual perceptual units a bird’s visual system would register, then running field predation trials with real moth targets pinned across ten unpolluted woodland sites in southern England [8]. Pale typica matched lichen backgrounds far more closely than carbonaria did by the avian-vision metric, and the field trial found pale moths surviving at a rate roughly twenty-one per cent higher than dark moths in that unpolluted setting [8]. That is the first time in the history of the case that “camouflage” stopped being a plausible-sounding word applied to a moth and a background, and became a number computed through a model of the actual predator’s eyes.

A small reflectance-measurement probe positioned just above a pinned typica specimen set against a lichened bark panel, the probe's tip still short of contact and a loupe left resting nearby
Figure 5. Bird-vision modelling later put a number on what Kettlewell only argued by eye: pale moths against lichen produced roughly twenty-one per cent higher survival than dark moths did, measured in the units a bird's own colour vision actually uses.Image prompt and art direction by Brecht Corbeel; generation pending.

The decline the mechanism predicted arrived on schedule, in two countries

A predation-based explanation for the rise of carbonaria makes an obvious further prediction that has nothing to do with defending the original result: if pollution declines, the camouflage advantage should reverse, and melanic frequency should fall, on a timescale set by the moth’s one-year generation time rather than by any human decision about how the story ought to end. Britain’s Clean Air Acts, beginning in 1956 and tightened through the following decades, gave that prediction a real natural experiment, running independently of anyone involved in the academic dispute.

The decline arrived, and it was tracked in enough detail to be more than an anecdote. Cook, Dennis and Mani’s 1999 survey of the Manchester area — the same district where the phenomenon had first been recorded a century and a half earlier — found melanic frequency at a peak of about ninety per cent in 1983, already falling to below ten per cent by the time of publication in 1999, and described the industrial melanism the species had lent its name to as, in their own words, now almost past in that district [4]. Cook’s 2003 review extends the same pattern nationally: the high-frequency plateau that had covered most of northeastern England contracted through the 1980s until it survived only in the far north, and by the early 2000s the maximum frequency recorded anywhere had dropped below fifty per cent, with most sites below ten per cent — a decline the paper models as requiring selection against the melanic form on the order of five to twenty per cent, the same range of magnitude, arrived at by entirely different methods, as the predation-based estimate from Majerus’s garden a few years later [5]. Cook and Saccheri’s 2012 retrospective adds two more independent readings of the same decline, both derived purely from tracking frequency change over time rather than from watching a single predation event: an Open University survey run in 1983 and 1984 implied an average disadvantage for the melanic form of about twelve per cent, and the long-running Rothamsted Insect Survey, covering 1974 to 1999, implied a disadvantage of about ten per cent [2]. None of these estimates were built to agree with each other — a garden predation experiment, a national frequency survey, and a standardized moth-trapping network are three different instruments pointed at three different kinds of data — and the fact that they cluster in the same five-to-twenty-per-cent band is the strongest kind of corroboration available in field biology: independent methods converging on the same number because the number describes something real in the population, not an artefact of how any one of them happened to be built.

The same reversal turned up independently on a different continent, driven by the same class of environmental cause. Grant and Wiseman’s 2002 survey of Michigan and Pennsylvania sites found melanic frequency collapsing from above ninety per cent in 1959 to about six per cent by 2001, tracking the American equivalent of Britain’s clean-air legislation and matching, point for point, the sulphur dioxide and particulate trends Grant, Owen and Clarke had documented running in parallel with the British decline in their earlier 1996 study [9, 10]. Two populations on two continents had risen together for the same apparent reason in the nineteenth century and, decades later, fallen together for the same apparent reason as well — a prediction nobody had designed the moth to satisfy, confirmed by a policy intervention aimed at human respiratory health rather than at settling an argument in evolutionary biology.

A modern moth trap's egg-box liner tipped out onto a collection-room bench, almost entirely pale typica moths with a single carbonaria specimen being lifted clear with forceps for logging
Figure 6. After Britain's clean-air legislation, carbonaria frequency fell from a plateau above ninety per cent to under ten per cent across most of its former range, and an independent decline of the same size showed up in Michigan and Pennsylvania on the same schedule.Image prompt and art direction by Brecht Corbeel; generation pending.

The self-correction is the demonstration, not the embarrassment

Laid end to end, the history runs: an observation of geographic correlation; an experiment that measured a real, direction-dependent survival difference; a legitimate methodological critique of that experiment’s realism; a public distortion of that critique into a much larger claim than its author intended or the evidence supported; a second, more careful experiment designed by the critic himself to test his own objection on its own terms; a posthumous, statistically explicit vindication of the original mechanism at a measured magnitude; independent molecular genetics that dated the causal mutation to a specific year consistent with the historical pollution record; a quantified, predator’s-eye measurement of the camouflage itself; and a decline, on two continents, that arrived exactly where the mechanism said it should once the environmental pressure driving it was removed.

None of that sequence required anyone to protect the case from scrutiny. Majerus’s critique was allowed to stand, was taken seriously by the community that had built the textbook example, and was answered with more data rather than with an appeal to authority — by the same person who raised it, at the cost of seven years of his own research life. Nobody suppressed the 1998 book, nobody prevented Coyne from publishing a sceptical review, and nobody stopped Hooper from publishing hers; the correction ran entirely through more publication, more fieldwork and more genetic sequencing, not through less of any of them. The creationist literature that treated a scientist’s provisional caution as a retraction and a journalist’s fraud allegation as an established finding was working from a much older and less flattering model of how science handles internal disagreement, one where a raised doubt is either suppressed or fatal. What actually happened is the opposite of both: the doubt was published, taken as a design specification for a better experiment, and the better experiment agreed with the original conclusion at a precisely stated magnitude, the number itself later corroborated from an entirely separate direction by the molecular dating of the mutation and by a direct measurement of avian-visible camouflage. A textbook case that can absorb its most serious critique, name the critique’s author as the person who ultimately settled it, and come out the other side with a mutation dated to a specific year is not a case that survived a scandal by accident. It is what the underlying claim — that natural selection is a real, measurable, falsifiable process rather than a story arranged after the fact — predicts should happen when the claim is actually true.

Measure the case by what would have falsified it at each stage, and the strength of the outcome becomes clearer still. Kettlewell’s design would have failed to show a reversal between Birmingham and Dorset if predation were morph-blind; it did not fail. Majerus’s garden would have failed to show selection against carbonaria at natural density and on natural resting sites if the original effect had been a pure artefact of unnatural placement; it did not fail either, and the resting-site data he collected specifically to test that possibility showed his own predicted artefact was not doing the work critics feared. Van 't Hof’s molecular dating would have landed on any of two centuries of possible dates if the transposition were unrelated to industrial pollution; instead it landed within a few decades of the documented start of heavy coal burning. Each stage offered a real chance for the mechanism to be wrong, and at each stage the data that came back were the data a true camouflage-and-predation explanation, and not a coincidence dressed up as one, should produce.