A Drought Turned Beak Depth Into a Matter of Survival
Daphne Major is a treeless volcanic cone in the middle of the Galapagos archipelago, a tuff crater whose rim rises roughly 120 metres above the sea, its floor and slopes crusted with dry scrub across a total area of about 0.34 square kilometres [10, 3]. In 1977 the island received almost no rain: the year’s total was 24 millimetres, and the resident finches did not breed at all [3]. What happened on that islet over the following months is the actual founding event of this article, not the anchoring of a survey ship four hundred kilometres away in 1835.
Peter Boag and Peter Grant had been catching, banding, measuring and releasing every Geospiza fortis, the medium ground finch, they could find on Daphne Major since 1973, logging beak length, beak depth, beak width, wing length, tarsus length and weight to fractions of a millimetre and of a gram on every bird [1, 2]. The mechanism behind what the drought then did to that population was already understood from the group’s own earlier fieldwork: a finch population’s mean beak size reflects a tradeoff between two kinds of seed. Small, soft seeds are cheap to handle but run out first in a long dry season; large, hard seeds remain abundant even in drought, because almost nothing else can eat them, but only a large enough beak can crack one open at all, and even a capable bird pays a real time cost per seed [3]. When Daphne Major’s small seeds were exhausted in 1977, that tradeoff briefly became the only one that mattered: a finch either had the beak to open what was left, or it did not eat.
Grant and Grant’s later tabulation of that year’s data puts numbers on what “intense” meant. Comparing the finches present before the drought with the finches that survived it, the average standardized selection differential — the difference in trait means between the two groups, expressed in units of the population’s own standard deviation — was 0.642 for males and 0.668 for females, and every one of the six measured traits moved the same direction: heavier weight, longer wings, longer and deeper and wider beaks, all favored, for both sexes, without exception [3]. Beak depth alone carried a differential of 0.80 in males and 0.69 in females; overall beak size, the first principal component of the three beak measurements combined, carried 0.80 and 0.74 [3]. Boag and Grant’s original 1981 report of the episode described the resulting selection intensities as the highest yet recorded for any wild vertebrate population, a description later work has not overturned [1].
Daphne Major’s usefulness as a research site comes from the same smallness that makes it inhospitable: no permanent human presence, no agriculture, and a population of both resident ground-finch species small enough in an average year that catching, marking and measuring essentially every bird on the island is a literal description of the fieldwork rather than a figure of speech [2]. None of what follows required inferring what natural selection must have done to explain a pattern left in rock or in a museum drawer. It was watched, on marked individual birds, as it happened, and then written down.
The Voyage Barely Noticed Them
The finches did not make Charles Darwin an evolutionist in 1835, and the record of what he actually did with them undercuts most retellings badly enough that the discrepancy is itself informative.
Frank Sulloway’s reconstruction of Darwin’s collecting practice aboard the Beagle, built from Darwin’s own ornithological notes and specimen labels, found that Darwin recorded island localities for only three of his thirty-one finch specimens [9]. That is not an oversight that happened before he had reason to be careful: Darwin had already noticed, and recorded, that the archipelago’s mockingbirds seemed to differ from island to island — notes on that pattern date to late September 1835, while he was still in the islands — and Sulloway’s reading of the specimen record is that Darwin did not change his finch-collecting or labelling habits afterward even so [9]. Darwin conceded the consequence in print later: writing in 1845, he admitted that most of his finch specimens “were mingled together” with no island data attached [9].
The bird that actually caught Darwin’s attention as a possible witness to transmutation during the voyage itself was the mockingbird, not the finch. The finches acquired their evidentiary weight only after the voyage, and only through someone else’s expertise: the ornithologist John Gould, working through Darwin’s specimens back in London, described what he judged to be twelve new finch species on 10 January 1837 and recognized that the extraordinary variation in their bills, across birds otherwise so similar, was the pattern worth explaining [9]. Darwin’s shift toward transmutation followed shortly after he met with Gould that March, once the geographic patterning Gould had identified in the finches was laid alongside what Darwin already suspected of the mockingbirds [9]. A specialist’s museum-bench comparison did the evidentiary work, eighteen months after the beach, not the beach itself.
The name “Darwin’s finches” is later still and belongs to the ornithologist David Lack, whose 1947 monograph of that title turned Gould’s taxonomic puzzle into the textbook case it is treated as now [3]. Scale alone makes the contrast with the modern program blunt. Darwin’s entire Beagle collection of Geospiza amounted to thirty-one specimens gathered across several islands in about five weeks ashore, most without so much as an island label attached [9]. The program this article is actually about has, since 1973, individually marked, measured, and in most years pedigreed nearly every bird of two resident species living on one 0.34-square-kilometre islet, continuously, for more than fifty years [2]. These are not two chapters of one research program. They are different orders of evidence entirely, and only one of them is capable of producing a selection differential, a heritability estimate, or a documented reversal.
Every Bird Got a Number, and the Number Made Selection Measurable
What separates the Daphne Major program from a single dramatic drought story is the accounting underneath it. Since 1973, the team has caught, individually marked, and repeatedly re-measured essentially every Geospiza fortis and G. scandens on the island every year, recording each bird’s parentage where behavioral observation of mating could establish it; the same six external measurements — weight, wing length, tarsus length, beak length, beak depth, beak width — were taken on the same protocol year after year, then reduced by principal components analysis into three interpretable synthetic traits: body size, beak size, and beak shape [2]. Annual sample sizes swung with population booms and droughts: in G. fortis alone, a single year’s catch ran from as few as 45 birds in 1997 to as many as 976 in 1991 [2].
The instrument doing the actual work is unglamorous: a small caliper, read to a fraction of a millimetre, closed on a living bird’s bill before the bird is released. The team checked, rather than assumed, that decades of caliper readings meant the same thing across time. In 2001 Peter Grant re-measured nine museum specimens of G. fortis that he had already measured in 1975 and 1976, and found no detectable difference among the three sets of readings on any of five traits [2]. That kind of methodological housekeeping is what allows a 1977 measurement and a 2005 measurement of the same trait, on different individuals thirty years apart, to be compared as though they came from one continuous instrument.
The measurement matters because a change in a trait’s mean value between two years is not automatically evolution. It could reflect which birds happened to be caught, or which happened to survive, in a population whose underlying genetics never moved. What licenses treating a selection differential as a claim about evolution is heritability: the degree to which offspring resemble their parents in the trait, quantified from the same pedigree data the program was already collecting. Peter Boag’s 1983 analysis of the Daphne Major population found high repeatabilities for the external measurements and high heritability for the derived beak and body traits in G. fortis in both years he examined [7]. The heritability estimates themselves needed correcting once genetic methods caught up with the behavioral pedigrees: some proportion of offspring in any wild bird population are not, despite appearances, the social father’s genetic offspring, since extra-pair copulations are common in Geospiza as in most passerines, and treating a chick as related to a male it merely shares a nest with distorts the apparent heritability of some traits. Correcting for misidentified paternity narrowed some estimates without erasing the pattern: by Grant and Grant’s 2002 summary, the corrected heritability of the morphological traits ran from about 0.5 to 0.9, high enough to make a quantitative prediction worth attempting [2].
Heritability is what turns a measured selection differential into a testable prediction, through the simplest quantitative-genetics identity in the field, the breeder’s equation:
where
The Direction Reversed Within Six Years
Selection on Daphne Major has never run in one direction for long, which is the detail most retellings drop when they use 1977 alone to illustrate natural selection in action. The trait that moved toward larger, deeper beaks in 1977 did not stay there.
Rain returned in 1978, and then, in 1982-83, an exceptionally strong El Nino event brought an extraordinarily prolonged wet season to the islands [2]. The composition of the seed supply flipped with it: small, soft seeds, scarce during the preceding drought, became abundant again, and large hard seeds grew comparatively scarce [3]. Gibbs and Grant, examining the finch population’s response to that swing, documented what they called oscillating selection: the same population that had favored large body size in the 1977 drought reversed course in the years following the opposite climatic extreme, on the same traits [6]. Beak shape, the second synthetic trait in the long-term dataset, moved abruptly toward a more pointed form in the mid-1980s and stayed there for the next fifteen years [2].
Selection on body and beak size traits was, in fact, common rather than rare across the full record: G. fortis experienced a statistically detectable selection episode in roughly one year in three, about once every 4.5 years relative to the species’ own generation time, and G. scandens at a comparable frequency relative to its 5.5-year generation time; beak shape was selected far less often than size in either species [2]. What made 1977 and 1983 worth naming individually, rather than folding into that background rate, was not that selection occurred — it occurred routinely — but that its magnitude in those episodes ran several times the study’s own median differential, which across the thirty-year record was typically only about 0.03 to 0.06 standard deviations and rarely exceeded 0.50 [2].
The 1983 wet season also did something the drought could not have: it made cross-species breeding survivable. Grant and Grant record that successful F1 hybrid offspring between G. fortis and the cactus finch G. scandens, and backcrosses of those hybrids into both parent populations, began being documented starting in that prolonged wet year, after having been essentially unrecorded before it [2]. Over the following decades that introgression, not selection, became the dominant force reshaping G. scandens: the variance in that species’ own beak-shape measurements doubled, from 0.430 in 1973 to 1.026 in 2001, almost entirely because of hybrid and backcross individuals entering the breeding population, while G. fortis’s beak-shape variance over the same span barely moved [2]. A population’s mean trait value is not moved by selection alone even in years when selection is measurably acting; gene flow between species that ordinarily do not interbreed can move it on its own, and on Daphne Major it did.
A Second Finch Changed What Fortis’s Beak Was For
In the virtual absence of the smallest ground finch, G. fuliginosa, Daphne Major’s resident G. fortis had for decades been unusually small in beak and body size for its species — the textbook example of what is called character release, a population’s tendency to expand into a competitor’s niche once that competitor is gone [3]. That situation changed in 1982, when a new competitor arrived to stay.
Occasional individuals of the large ground finch, G. magnirostris — at roughly 30 grams, nearly twice the mass of an 18-gram G. fortis and well over twice a 12-gram G. fuliginosa — had visited Daphne Major in the dry season since at least 1973 without ever breeding there [3]. In late 1982, at the start of the same exceptionally strong El Nino that brought 1,359 millimetres of rain to the island that year, two females and three males established the first breeding population [3]. G. magnirostris is built to specialize on one difficult resource: the seed of Tribulus cistoides, held inside a hard, woody mericarp that has to be cracked or torn open before the seed inside is reachable [3]. G. fortis’s largest-beaked individuals — the same birds that had disproportionately survived the 1977 drought — can open a Tribulus mericarp too, but on average take three times as long to do it as a G. magnirostris does, and the smallest members of the fortis population cannot manage it at all [3].
For two decades the species’ numbers stayed too far apart for the overlap to matter much. G. magnirostris grew slowly through local breeding and immigration, reaching a peak of 354 ± 47 birds in 2003 [3]. Then the rain failed again: sixteen millimetres in 2003, twenty-five in 2004, no breeding in either year, and both species’ numbers collapsed under a food supply neither could count on being replenished [3]. What made 2004 different from 1977 was not the rainfall, which was almost identical, but the presence of a resident competitor at high density: roughly 150 G. magnirostris against roughly 235 G. fortis at the start of the year, with the two species’ total biomass close to equal, because each magnirostris weighed about twice what a fortis did [3]. Direct feeding observations from that year show the squeeze: a minimum of ninety birds observed foraging on Tribulus mericarps for 200 to 300 seconds apiece failed to extract seeds from more than two mericarps each, whereas in the 1970s, before magnirostris was established, eight birds observed for the same span worked through nine to twenty-two mericarps apiece, opening a fresh one on average every 5.5 seconds [3]. The large seeds that had saved the large-beaked survivors of 1977 were, by 2004, being consumed by a competitor that could reach them faster, and G. fortis’s own recorded use of Tribulus in its diet fell from 16.7 percent of feeding observations in 1977 to 8.2 percent in 2004, a statistically significant drop [3].
The Strongest Response the Study Ever Recorded
From 2004 into 2005, with both species starving on a seed supply too depleted to support either, G. fortis experienced the strongest directional selection episode Grant and Grant had measured in thirty-three years of the study — and this time the direction was reversed from 1977: it punished large beaks rather than rewarding them [3]. Selection differentials on the six measured traits were uniformly negative in both sexes, averaging 0.774 standard deviations in males and 0.649 in females; beak length alone carried a differential of −1.08 in males and −0.95 in females, and overall beak size, the first principal component, −1.02 and −0.92 [3]. Because beak depth and beak width are so tightly correlated in this population (r = 0.861 in males, 0.946 in females) that selection cannot statistically separate their independent effects, a selection-gradient analysis identified beak length as the one trait entering the model as an independent, statistically significant predictor of survival on its own, in both sexes [3].
G. magnirostris, meanwhile, was not selectively dying at all: its own four surviving males did not differ from thirty-two non-survivors on any of the six measured traits, and only one of thirty-eight measured females survived the crash — the competitor species was simply being starved out wholesale, not filtered by beak morphology [3]. By 2005 the G. fortis population had fallen to 83 individuals, its lowest count since the study began in 1973, and G. magnirostris to four females and nine males. Of the birds that disappeared, 13.0 percent of the missing magnirostris and 21.7 percent of the missing fortis were found dead, and every one of those carcasses — 23 magnirostris, 45 fortis — had an empty stomach [3].
The evolutionary response predicted from this selection episode is checkable, and Grant and Grant checked it. The mean beak size of the 2005 generation, measured the following year, was significantly smaller than the mean of the 2004 parental generation before selection had acted (t = 4.844, P < 0.0001), a shift of 0.70 standard deviations [3]. Run back through the breeder’s equation from the previous section, using the measured selection differentials and the confidence interval on the heritability estimate, the predicted range for that response was 0.66 to 1.00 standard deviations, and the observed 0.70 fell inside it [3]. Grant and Grant’s own summary of the result leaves little to interpretation: this was, in their words, the strongest evolutionary change seen in the thirty-three years of the study [3]. It is also, by the criterion set out decades earlier for character displacement — divergence in a resource-exploiting trait, caused specifically by competition, documented from the competitor’s arrival through to the measured evolutionary change — one of the only times that full causal chain has been recorded directly in a wild population, rather than inferred from a snapshot of two established species already sitting at different sizes [3].
Two Genes Now Have Names
Everything described so far is phenotype and pedigree: beaks measured with calipers, parentage inferred from watching who mated with whom. A separate, independent line of evidence arrived a decade later from the birds’ DNA, and it converges on the same events rather than merely restating them.
In 2015, a team led by Sangeet Lamichhaney and Leif Andersson published whole-genome sequences from 120 birds spanning every recognized species of Darwin’s finch plus two closely related mainland species, sampling individuals across as many as six different islands per species [4]. Comparing genomes across that sample, they identified a roughly 240-kilobase stretch of DNA containing the gene ALX1 — a transcription factor already known to be essential for craniofacial development in vertebrates generally — whose variants tracked with differences in beak shape both between species and within the single species G. fortis itself [4]. The same genomic survey found that gene flow between species, ordinarily assumed rare once lineages have split, had instead been a persistent feature of the whole radiation’s history, including a hybridization event between the ancestral warbler-finch lineage and the common ancestor of the tree and ground finches roughly a million years ago [4]. That gene flow is not a side note to the Daphne Major story; it is the same mechanism, on a longer timescale, that Grant and Grant had already caught happening in real time in G. scandens after 1983.
Beak shape was ALX1’s signature. Beak size, the trait that actually moved in 1977 and again in 2004-05, had its own locus, identified the following year by the same research group, this time focused specifically on the drought episode itself. Lamichhaney and colleagues sequenced the genomes of finches sampled across the size range on Daphne Major and identified a roughly 525-kilobase region containing the gene HMGA2 — a gene already known, in dogs, horses and humans, as one of the most consistent genetic correlates of body size and stature — as the strongest candidate locus for beak-size variation in the medium ground finch [5]. Genotyping a diagnostic marker within that region across 133 individuals, they found it accounted for roughly 27 percent of the phenotypic variance in beak size within the Daphne Major G. fortis population on its own [5]. And when they looked specifically at survival through the 2004-05 drought, the genotype associated with large beak size carried a selection coefficient of 0.59 against it — a strong selective disadvantage, measured not from morphology but directly from which genetic variant a bird carried into the crash [5].
Neither locus is the whole explanation, and the papers do not claim it is. HMGA2’s 27 percent of variance in one population, in one drought, leaves most of the remaining variation in beak size attributable to other loci, to environmental effects on growth, or to both; the same research program describes beak variation across the finch radiation more broadly as the product of many gene variants of individually small effect, with ALX1 and HMGA2 standing out mainly because their effects happened to be large enough to detect at all in a population this size [5]. The genomic work narrows the mechanism. It does not replace the decades of phenotypic measurement that first showed there was a mechanism worth looking for.
The ALX1 paper’s publication date, 12 February 2015, fell one day before what would have been Darwin’s two hundred and sixth birthday [4]. The coincidence is trivial; what is not trivial is that neither paper needed Darwin’s own reasoning to reach its conclusion. Boag, Grant and Grant established the character-displacement event from banding records, calipers and pedigrees alone, years before anyone had sequenced a single finch genome. The genomic work did not discover that event. It explained, at the level of a specific stretch of DNA, why the trait the calipers had been tracking for thirty years was there to be selected on at all.
Forty Years Made “Unpredictable” the Honest Word
By 2001, Grant and Grant had thirty years of continuous annual data on G. fortis and the cactus finch G. scandens, and they used it to ask a question no snapshot study could answer: not whether natural selection happens, but whether its long-run outcome can be predicted from what is already known about selection and heredity in the short run. Their answer, stated in the title of the paper reporting it, was that evolution can be predicted in the short term from a knowledge of selection and inheritance, but that in the long term it is unpredictable, because the environments that set the direction and strength of selection themselves fluctuate unpredictably [2].
The record backs the claim on both halves. Over those thirty years, G. fortis underwent eight statistically distinct evolutionary events — four in body size, three in beak size, one in beak shape — and G. scandens seven, two in body size and five in beak size, with selection sometimes unidirectional for years running and at other times reversing outright [2]. Each individual response, immediately following its own selection episode, was close to what the breeder’s equation predicted from the measured differential and the trait’s known heritability; that is the short-term predictability. But strung together over three decades, the sequence added up to something neither Grant could have forecast from the 1973 starting values. The phenotypic states of both species at the end of the thirty-year study, in the authors’ own words, could not have been predicted at the beginning, and they note plainly that if they had stopped sampling after only ten years, their conclusions about the direction of change in G. fortis beak size specifically would have been wrong, because the population’s mean had not yet turned back from where it stood in the early 1980s [2].
Grant and Grant’s own closing argument in that 2002 paper reached beyond their two study species: field studies like theirs, combined with multigenerational laboratory studies of microorganisms and experimental field manipulations of selection in other taxa, were, in their view, the only reliable way to build a bridge from measured microevolution on the scale of decades to the speciation and further adaptive radiation that operate on the scale of hundreds of thousands of years [2]. Whether any one generation’s measured shift in beak size ever accumulates into a new species is not a question this dataset alone can answer; it is a question about what happens when enough droughts and enough arriving competitors line up, in the same direction, for long enough that reversal stops being the likely outcome.
The 2004-05 character-displacement event, arriving three years after that 2002 paper, was itself unpredicted by the model then in hand: nothing in the thirty-year record up to 2001 involved a second competing species changing what G. fortis’s own beak was being selected for. The Grants’ 2014 monograph, 40 Years of Evolution: Darwin’s Finches on Daphne Major Island, folded that event and the decade following it into the continuous record, and a revised edition has since added the genomic findings on ALX1 and HMGA2 to the same account, so that the same specimens tracked by caliper since 1973 are now also tracked by genotype [8].
What more than four decades of that arrangement changed is not whether natural selection is real — Boag and Grant’s 1981 paper already put that beyond reasonable dispute for anyone willing to read a table of selection differentials — but what kind of claim “natural selection acted here” is allowed to be. It is no longer an inference run backward from a pattern of resemblance among living and fossil forms, the kind of evidence Darwin actually had in 1859 and, for that matter, in 1835. On Daphne Major it is a measurement, repeated on marked individuals, checked against a quantitative prediction, and reversed in direction at least twice within the observers’ own working lives. The honest description of what that measurement shows is not that evolution is directional, or progressive, or even reliably forecastable past a single generation. It is that selection is real, frequent, and strong enough to move a wild population’s mean trait values by most of a standard deviation within a single drought, and that which way it will move next is a question about next year’s rainfall and this year’s competitors, not about beaks.