A diagram outlived the model it was drawn for

Almost every biologist can draw the adaptive landscape from memory: a range of hills, a population as a ball or a cloud of dots somewhere on a slope, selection pushing it uphill, valleys of low fitness separating one summit from a higher one nearby. The picture is nearly a century old, it is genuinely useful, and it has been doing a great deal of unlicensed intellectual work the entire time.

The trouble is that two distinct objects travel under the same name. One is a formal map: an assignment of a fitness value to every possible genetic constitution. That object is well defined, it can be measured for small systems, and it supports theorems. The other is a rendering of that map as a two-dimensional surface embedded in three-dimensional space, complete with slopes, ridges, saddles and basins. The rendering is a projection, and projections lose things. What this one loses is precisely the property that matters most for how evolution actually moves: the dimensionality of the space being walked across.

This article separates the two. It reconstructs what Sewall Wright proposed and why, states the dispute with R. A. Fisher without pretending it has been settled, and then asks how much of the hills-and-valleys intuition survives contact with high-dimensional genotype space, with epistasis, with combinatorially complete laboratory reconstructions, and with the fact that the surface is deformed by the population standing on it.

ADVERTISEMENT

What Wright actually proposed

Wright’s mathematical work on the joint action of mutation, selection, migration, inbreeding and random drift in subdivided populations predates the famous diagram [1]. In that framework, the fate of a gene combination is not determined by selection alone; it is determined by the relative magnitudes of selection, mutation pressure, migration between local demes, and sampling error, which is itself a function of effective population size. The landscape diagram, presented at the Sixth International Congress of Genetics in 1932, was a visual summary of that argument, not a new theory.

Wright’s proposed resolution — the shifting balance process — is conventionally described in three phases: random drift in small, partially isolated local populations carries some deme into a new domain of attraction; mass selection within that deme then drives it up the new peak; and finally differential dispersal or interdemic selection spreads the superior combination across the species. Wright’s point, as it is usually reconstructed by later critics and defenders alike, was that selection acting alone on a large panmictic population is a hill-climbing procedure and hill-climbing procedures get stuck [2, 3].

Two features of the original are worth holding onto, because both are routinely dropped. First, the diagram was always understood to be a compression of something far larger; Wright was drawing a two-dimensional slice through a space of gene combinations, not asserting the space had two dimensions. Second, the mechanism is fundamentally about population structure. A shifting balance argument without subdivision is not a shifting balance argument at all.

A winter coppice wood worked in separate panels, the nearest one caught half cut with low fresh stools on one side of the cutting line and tall uncut rods still standing on the other, and further panels beyond at plainly different heights of regrowth
Figure 1. Shifting balance is an argument about subdivision; without separate parcels running out of step with one another there is nothing for selection between them to act on.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

The dispute with Fisher is not closed

Fisher’s position was that adaptive evolution in large populations does not need drift to escape anything, because in a sufficiently large and variable population the relevant environment and genetic background are changing continuously and mass selection suffices. This is a real disagreement about the architecture of adaptation, not a misunderstanding, and it has never been resolved by a decisive experiment.

The most influential modern statement of the sceptical case concluded that phases of the shifting balance process can occur under some conditions, but that genetic drift is often unnecessary for movement between peaks, that the conditions required for the third phase to spread a new combination through a species are restrictive, and that there was little empirical evidence that the process is more effective than simple mass selection [2]. That paper is a critique, and it should be read as one: it argues that the burden of proof has not been met, not that the process is impossible.

ADVERTISEMENT

The counter-case is equally serious and equally live. Reviewing the same theories in the light of metapopulation ecology, Wade and Goodnight argued that local extinction and recolonisation are common in natural populations and have genetic consequences neither Wright nor Fisher fully modelled, and that drift in small demes can convert non-additive, epistatic variance into additive variance — thereby increasing rather than limiting the potential for local adaptation [3]. On this reading, the disagreement is less about whether drift can matter than about which idealisation, the large panmictic population or the structured metapopulation, is the better default for real species.

The honest summary is that the field has not converged. Both camps agree on the underlying population genetics; they disagree about which parameter regimes are typical in nature, and typicality is an empirical claim about a domain that is extremely hard to sample. Anyone who tells you the dispute was won is reporting a preference, not a result.

The axes were never what the picture implies

Here is the first place the diagram actively misleads, and it is a subtle one. Two incompatible readings of the horizontal axes circulate. In one, the axes are allele or genotype frequencies, so a point is a population and the surface height is that population’s mean fitness. In the other, the axes are genotypes themselves, so a point is an individual genetic constitution and the height is its fitness. These are different spaces with different topologies, and results proved in one do not transfer automatically to the other. Much of the confusion around “peaks” comes from sliding between them.

Take the genotype-space reading, which is the one that empirical work now uses. A genotype is a sequence. The map is

w ⁣:GR,G=kL w \colon \mathcal{G} \to \mathbb{R}, \qquad |\mathcal{G}| = k^{L}

where the sequence has some length and each site takes one of a small number of states. Two genotypes are neighbours when they differ at a single site. This is the object Maynard Smith described as protein space, and his framing remains the cleanest statement of the constraint: evolution must move through the space one step at a time, and every intermediate must itself be viable, in the way that changing WORD to GENE one letter at a time requires each intervening string to be a real word [4].

Now notice what the picture cannot show. A sequence of even modest length has a number of one-step neighbours equal to its length times the number of alternative states per site — dozens or hundreds, not the four compass directions a surface allows. Kauffman and Levin’s analysis of adaptive walks on rugged landscapes is explicitly built on this neighbourhood structure rather than on any surface geometry, and one of its central results is that as the complexity of the entities under selection increases, the local optima reachable by such walks fall progressively closer to the mean fitness of the space [6]. That is a statement about combinatorics, and it has no natural depiction as a hill.

ADVERTISEMENT
A single hazel stool seen from low down looking up, dozens of straight rods springing from one low stump and fanning outward in every direction against a bright pale sky, one severed rod hanging clear of the cut and caught among its neighbours
Figure 2. A single state has dozens of one-step neighbours, not the four directions a drawn surface allows; the count is combinatorial, and no hill can display it.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Gavrilets pressed the point further and turned it into an argument against the metaphor itself. Because the dimension of real adaptive landscapes is vastly larger than three, the intuitions imported from three-dimensional experience are unreliable; in high-dimensional genotype spaces, sufficiently fit genotypes can form a connected, percolating set threading through the space, so that a population can move between distant high-fitness regions without ever descending through a valley. On such a “holey” landscape, the problem Wright’s shifting balance was designed to solve may simply not arise [5]. This is not a refutation of Wright’s population genetics. It is a claim that the geometric framing overstated the difficulty.

Epistasis makes the surface rugged; sign epistasis is what actually blocks a path

Ruggedness in a landscape means multiple local optima, and its source is epistasis: the fitness effect of a mutation depending on the genetic background it lands in. If every mutation had a fixed effect regardless of background, the landscape would have a single peak and every uphill path would reach it. Ruggedness is epistasis made geometric.

But not all epistasis constrains paths. The consequential form is narrower. Weinreich, Watson and Chao named it sign epistasis: the case in which the sign, not merely the magnitude, of a mutation’s fitness effect is under epistatic control, so the same mutation is beneficial on some backgrounds and deleterious on others [7]. This distinction is the hinge of the whole subject. Magnitude epistasis changes how fast a population climbs. Sign epistasis changes what it can climb at all.

The reason is an ordering constraint. Under the standard assumption that a population moves by fixing one beneficial mutation at a time, a mutational path is accessible only if fitness increases at every single step:

w(g0)<w(g1)<<w(gL) w(g_0) < w(g_1) < \cdots < w(g_L)

A single step that violates this inequality closes the entire path, regardless of how high the endpoint is. Sign epistasis is exactly the mechanism that produces such steps, and reciprocal sign epistasis — where each of two mutations is deleterious without the other and beneficial with it — is what creates genuinely separated local optima.

The assumption embedded in that inequality deserves to be made explicit rather than absorbed. It is the strong-selection, weak-mutation regime, in which the population is effectively monomorphic between substitutions. It is a modelling choice with a limited domain, not a law, and much of the interesting behaviour discussed below consists of what happens when it fails.

A line of driven hazel stakes along a laid hedge with a long springy binder rod caught mid-weave in the air, woven behind stakes on one side and bare stakes waiting on the other
Figure 3. Path accessibility is an ordering constraint, not a distance: a binder can only be woven where the stakes already stand.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

What the combinatorial reconstructions showed, and what they did not

The subject stopped being purely theoretical when it became possible to construct every combination of a small set of mutations and measure each one. The landmark case examined five point mutations in a beta-lactamase allele that jointly increase bacterial resistance to a clinically important antibiotic by a factor of roughly one hundred thousand. In principle the five mutations can be acquired in any of 120 orders. Weinreich and colleagues found that 102 of those 120 trajectories are inaccessible to Darwinian selection, and that many of the remainder have negligible probability, because four of the five mutations fail to increase resistance in some combinations. They attributed the pattern to pervasive biophysical pleiotropy within the protein and concluded that much protein evolution will be similarly constrained, and therefore that the protein tape of life may be largely reproducible and even predictable [8].

That last inference is the contested part, and it should be flagged as an interpretation rather than a measurement. The measurement is the accessibility count. The extrapolation from one enzyme under one antibiotic to protein evolution in general is an argument, and reasonable people have resisted it.

Related work reconstructing intermediate forms in the laboratory made the complementary point that empirical landscapes let one finally ask why particular paths are taken rather than only observing the endpoints, and located the answer in epistatic interactions among the mutations involved [9]. Surveying the field a decade later, Hartl summarised two recurring findings from studies that construct all combinations at a set of sites: epistasis between beneficial mutations frequently shows diminishing returns, with favourable mutations performing worse in combination than their individual effects predict, and only a small fraction of the theoretically possible routes are accessible, largely because sign epistasis requires steps that temporarily decrease fitness [11].

Two caveats about this literature are structural rather than incidental, and de Visser and Krug set them out in their review of empirical landscapes and predictability [10]. First, the mutations chosen for these constructions are not random: they are typically drawn from phylogenies or from evolution experiments, which means they were selected precisely because they occurred together, biasing the sample toward interacting sets. Second, a combinatorially complete landscape over five sites is a tiny window cut into a space of astronomically many sequences, and the ruggedness of the window is not automatically the ruggedness of the whole.

Larger reconstructions have complicated the picture in both directions. A mutational scan of the green fluorescent protein across tens of thousands of derivative genotypes found the local fitness peak to be narrow, shaped by a high prevalence of epistatic interactions, with fluorescence lost once the joint effect of accumulated mutations crossed a threshold [12]. That result supports ruggedness. Pulling the other way, a genome-edited map of more than 260,000 genotypes of dihydrofolate reductase under trimethoprim found a landscape that is highly rugged — 514 fitness peaks — and yet whose highest peaks are reachable by abundant fitness-increasing paths, with overlapping domains of attraction that make the outcome contingent on chance rather than blocked [13]. Ruggedness and inaccessibility, it turns out, are not the same property, and the mountain-range picture conflates them by construction.

The ground is being cut while you stand on it

Every formulation above treats the map from genotype to fitness as fixed. It is not, and the metaphor’s static geology is a substantial part of why the assumption goes unexamined.

Fitness is a property of an organism in an environment, and for most organisms the environment includes conspecifics, competitors, hosts, parasites and predators that are themselves evolving. Where fitness depends on the frequencies of genotypes in the population, the surface is deformed by the very movement of the population across it: a genotype that is advantageous while rare loses that advantage as it becomes common. Under coevolution the deformation is driven from outside, by another lineage adapting in response. In neither case is there a stable summit to reach.

Mustonen and Lässig proposed replacing landscapes with seascapes for exactly this reason, treating adaptation as a non-equilibrium phenomenon — the genomic response to time-dependent selection — rather than as relaxation toward a fixed optimum, and arguing that adaptive evolution is signalled by a sustained surplus of beneficial over deleterious substitutions rather than by the presence of positive selection as such [14]. The reframing matters practically: the equilibrium reasoning that a static landscape invites can misclassify a population that is running to stay in place.

A freshly cut hazel stool with pale bright cut faces low in a coppice clearing, already overhung by taller multi-stem regrowth from an earlier cut, with a billhook set in a stump
Figure 4. The surface deforms as the population moves on it; a stool is never cut back into the same woodland it was cut from.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Drift, population size, and the networks that make valleys optional

If some paths really are blocked to strict hill-climbing, how does anything ever cross? The answer is that the strict hill-climbing assumption is the thing that breaks first.

A finite population does not deterministically fix only beneficial mutations. Deleterious or neutral intermediates can drift to appreciable frequency, and the rate at which a population traverses a valley depends jointly on population size, mutation rates and the fitness cost of the intermediates. Weissman and colleagues showed that the mechanism itself changes with population size: in small populations, intermediates drift to fixation one at a time in a sequential-fixation regime, while in large populations intermediates remain rare and a lineage tunnels through by acquiring the second mutation before the first has fixed. Their analysis found that valley crossing can be remarkably fast when the intermediates are close to neutral [15]. Note the direction of the population-size effect: it is not simply that small populations drift more and therefore cross better. Which regime applies, and which is faster, depends on the parameters.

The other resolution is that many valleys are not valleys at all but detours across flat ground. Where many sequences share a phenotype, they form neutral networks in genotype space through which a population can move without fitness cost. Wagner’s analysis of RNA structures used this to resolve the apparent paradox between robustness and evolvability by separating two levels: a genotype that is individually robust to mutation explores less, but a phenotype realised by a large connected set of sequences allows a population to spread across that set, accumulate diversity, and thereby reach a far greater range of neighbouring phenotypes [16]. This is the mechanism underlying the holey-landscape argument, and it is the strongest single reason to distrust the valley imagery: in a space of very high dimension, the neutral connections are numerous enough that a barrier in a two-dimensional cross-section is often no barrier at all.

A low rounded gap worn clean through the base of a thick laid hedge, its edges rubbed pale and smooth, with flattened frosted grass leading into it and bright open field showing through
Figure 5. A barrier in cross-section is often no barrier at all; the way through runs round and beneath, at a level the drawing never shows.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Where the metaphor earns its keep, and where it lies

An honest verdict has to be split, because the metaphor is neither empty nor safe.

It helps in four places. It makes the concept of multiple optima intelligible to anyone in about ten seconds. It correctly conveys that the fitness consequence of a mutation is a function of the whole genotype, which is the single most important corrective to thinking about genes one at a time. It gives experimentalists a target: the combinatorially complete studies above exist because the landscape framing made “measure the whole neighbourhood” an obvious thing to attempt. And it supplies a shared vocabulary — accessible path, local optimum, ruggedness — that has proved formalisable rather than merely evocative.

It misleads in four corresponding places, and each failure traces back to the same source, which is the substitution of three-dimensional geometry for combinatorial structure. It implies a small number of directions from any point when there are many. It implies that separation implies isolation, when high-dimensional connectivity and neutral networks routinely provide routes around an apparent barrier [5, 16]. It implies that ruggedness entails inaccessibility, which the dihydrofolate reductase map directly contradicts by exhibiting hundreds of peaks and abundant accessible paths at the same time [13]. And it implies a static surface, which frequency dependence and coevolution rule out for most organisms [14].

There is a philosophical position, which should be attributed rather than asserted, that the right response is to abandon the picture altogether and reason only from the formal models. The weaker position, which the empirical literature seems to support better, is that the picture is safe wherever it functions as an index to a formal object and unsafe wherever its geometry is doing the inference. The practical test is easy to apply: if an argument would change when the number of dimensions changes, the argument is being made by the drawing rather than by the model.

A prediction, with the condition that would falsify it

The following is a forecast, offered as such and separated from the findings above.

Prediction, horizon 2031. Directly measured landscapes will keep enlarging, and the modal published result will continue to be “rugged but navigable” rather than “rugged and blocked” — that is, most newly published combinatorially complete or genome-edited landscapes spanning more than one thousand genotypes will report multiple local optima together with a substantial number of accessible fitness-increasing routes to the highest of them.

Assumptions. That genome editing and deep mutational scanning costs continue to fall; that the systems chosen remain dominated by microbial enzymes and fluorescent reporters under laboratory selection; and that “accessible” continues to be operationalised through the monotone-increase criterion above.

Observable indicators. The ratio of reported accessible to inaccessible direct paths in new large landscapes; whether papers reporting many peaks also report reachability of the global optimum; and whether the field’s reviews shift their framing from constraint toward contingency.

Disconfirmation. The prediction fails if, within that horizon, several independent large-scale landscapes in different systems report that the global optimum is unreachable by any monotone path from typical starting genotypes, and that finding survives the obvious methodological objections about measurement noise and the choice of mutation set. It would also fail if the field abandons the monotone-path criterion so thoroughly that the accessible-path statistic stops being reported at all, in which case the prediction becomes unmeasurable rather than wrong.

The map that is not the territory, and not the picture either

Wright’s contribution was a set of equations about the interaction of selection, mutation, migration and drift in structured populations, and those equations remain load-bearing [1]. The diagram he drew to explain them has outlived the specific argument it illustrated and acquired an authority the argument never claimed. The empirical programme that grew out of the landscape idea has been genuinely productive, and it has repaid the metaphor by falsifying several of its visual implications [10, 11].

A laid hedge is a better mental image than a mountain range, if one insists on an image at all. The stems available to bend next year are entirely determined by which were cut and which were left this year, the geometry is local rather than panoramic, the same cut is right on one stem and fatal on another, and the whole boundary keeps growing while the work proceeds. Nothing about it invites the question of how to get from one summit to another, because there are no summits — only a state, its neighbours, and which of them can be reached from here.