A sequence is a specification, not a shape
The observation that started the field is narrow and robust. Anfinsen and colleagues showed that ribonuclease A, reduced and denatured so that its disulfide bonds were broken and its fold destroyed, recovered essentially full enzymatic activity when returned to permissive conditions and allowed to reoxidize, without any cellular machinery present [1]. The chain found its way back.
That is the fact. The interpretation built on it — the thermodynamic hypothesis — is that the native conformation of a protein is determined by its amino acid sequence, in a given environment, and corresponds to the state in which the free energy of the whole system is lowest. It is worth reading the qualifiers rather than skipping them. The claim is about a system, not a string: solvent, ionic conditions, temperature and redox state are all inside the boundary. And it is a claim about determination, not about attainability. It does not say that every chain reaches its global minimum inside a cell, where cotranslational folding, chaperones, kinetic traps and aggregation are all in play. Nor does it say anything at all about how a chain gets there, and still less about how an outside observer could compute the answer.
Those three questions — what the code is, how folding can be fast, and whether a machine can predict structure from sequence — were laid out as separable by Dill and MacCallum in their fiftieth-anniversary survey of the problem [4]. Keeping them separate is the whole discipline of this article. The last decade produced a striking answer to the third. It did much less to the first.
Levinthal’s arithmetic was a reductio, not a measurement
The standard framing of why search cannot be the mechanism is combinatorial. Treat each residue as having some small number of accessible backbone conformations, and the number of chain configurations is exponential in length:
Take a modest three states per residue and a hundred-residue chain. The arithmetic gives
Two things about this argument are routinely misreported. First, the specific numbers are illustrative and depend entirely on the assumed
Zwanzig, Szabo and Bagchi made the escape route quantitative: a small and physically reasonable energy bias against locally unfavourable configurations is sufficient to reduce the folding time from cosmological to biologically reasonable values [2]. The lesson is that the landscape need not be flat, and once it is not flat, the counting argument loses its force immediately. A very modest tilt buys an enormous amount.
The funnel is a claim about the statistics of a landscape
The systematic version of that tilt is the funnelled energy landscape. Bryngelson, Onuchic, Socci and Wolynes synthesised a statistical treatment of the energetics of protein conformation in which sequences that fold well are those whose landscapes are minimally frustrated: the energetic bias toward native-like contacts outweighs the ruggedness arising from competing non-native interactions [3]. On such a landscape, thermal motion is enough. Random fluctuations lead energetically downhill, so folding proceeds by a funnel of many descending routes rather than a single pathway threaded through an otherwise flat space [4].
Two clarifications keep this honest. The funnel drawing is a two-dimensional cartoon of an object with thousands of dimensions; the real content is a statistical statement about the relative magnitude of bias and ruggedness, not a literal smooth cone. And the framework explains why folding is possible and fast. It does not, by itself, tell anyone what a given sequence folds into. Physics-based simulation of folding remained, for decades, a demonstration achievable for small fast-folding domains under heroic computational effort rather than a general predictive tool [4].
The usable information turned out to be evolutionary
The breakthrough in prediction did not come from better force fields. It came from noticing that the answer had already been written down, many times over, in sequence databases.
The reasoning is straightforward. If two residues are in contact in the folded structure, they are structurally coupled: a substitution at one position that would disrupt the packing can be tolerated if it is compensated at the other. Across a family of homologous sequences that have all been folding successfully for hundreds of millions of years, that compensation leaves a statistical trace. The obstacle is that raw correlation between alignment columns is badly confounded. If A is coupled to B and B to C, A and C appear correlated without touching, and phylogenetic relatedness among the sampled sequences manufactures further spurious correlation.
The fix is to fit a global model that explains the observed column statistics with the smallest number of direct couplings. In the maximum-entropy formulation used in this literature, the probability of a full sequence takes a Potts form,
with single-site fields
This is the pivot on which everything after it turns, and it deserves to be stated plainly rather than absorbed into a general story about deep learning. The information source that made computational structure prediction work is the evolutionary record. A contact map inferred this way is not derived from the physics of the chain. It is read off a very large sample of chains that already solved the problem. Structure prediction of this kind is therefore inference over solved instances, and its reliability is expected to track the depth and diversity of the alignment available for a query — a dependence that later architectures have reduced but which has not been shown to vanish.
Note also the shape of what coevolution delivers: a two-dimensional map of pairwise proximities from which three-dimensional coordinates must still be reconstructed. That reconstruction is exactly the lofting problem — a lower-dimensional specification faired up into a full-size form.
CASP made the claim falsifiable before it made it famous
The methodological substrate for all of this is easy to overlook. In 1995 Moult, Pedersen, Judson and Fidelis described a large-scale experiment to assess protein structure prediction methods [7]. The design has barely changed since. Targets are sequences whose experimental structures have been determined but not yet released. Predictors submit models before the answers are public. Independent assessors, who did not build the methods, evaluate the submissions.
Every clause there is doing work. The test set is held out by construction, not by an author’s promise. It is chosen by experimentalists and organisers rather than by the people being graded, so it cannot be curated toward a method’s strengths. Submission precedes disclosure, which forecloses tuning on the answer. And the evaluation is performed by a third party, which removes the most common failure mode in computational biology: a paper reporting its own benchmark, on its own split, with its own metric.
The result of CASP14 in 2020 is best quoted from the assessors rather than from the developers. Kryshtafovych, Schwede, Topf, Fidelis and Moult reported that deep-learning methods from one group consistently delivered computed structures rivalling the corresponding experimental ones in accuracy, and described this as a solution to the classical protein-folding problem, at least for single proteins [8]. The trailing clause is theirs, not an editorial hedge added here.
The method behind those submissions was described by Jumper and colleagues at DeepMind, who reported a median backbone accuracy of 0.96 Å r.m.s.d.
What the blind assessment certified is specific: on CASP14’s target set, submitted models matched experimental structures within roughly the range of experimental disagreement for a majority of domains. It certified nothing about targets unlike CASP targets, and nothing whatever about quantities that are not a single static structure.
The confidence metrics are the most useful thing in the output
Structural comparison metrics matter here more than they usually do, because they determine what “accurate” means. Mariani, Biasini, Barbato and Schwede introduced lDDT, a superposition-free local score that evaluates distance differences over all atoms and is therefore insensitive to domain motions that would wreck a global superposition [11]. AlphaFold’s per-residue pLDDT is a prediction of that quantity: the model’s own estimate of the lDDT-Cα accuracy of each residue in its output [9]. The predicted aligned error, by contrast, estimates the expected positional error of one residue given alignment on another, and so speaks to relative domain placement rather than local quality — a distinction that Agarwal and McShan argue is regularly conflated in practice [21].
The proteome-scale deployment gives a sense of realistic yield. Tunyasuvunakool and colleagues reported predictions covering the large majority of the human proteome, with 58% of residues carrying a confident prediction, of which a subset amounting to 36% of all residues reached very high confidence [12]. Varadi and colleagues at EMBL-EBI then built the AlphaFold Protein Structure Database around such predictions, expanding structural coverage of sequence space by orders of magnitude relative to experimentally determined structures [13]. Read the human-proteome figure carefully: in the flagship application, rather fewer than two residues in five were very-high confidence.
Self-assessment also has limits that only external comparison exposes. Terwilliger and colleagues, comparing predictions against experimental maps, found many cases of remarkably close agreement but also cases in which even very-high-confidence predictions differed from the experimental data globally, through distortion and domain orientation, and locally, in backbone and side-chain conformation [14].
What a predicted static structure does not contain
An ensemble
A protein in solution is not a coordinate file. It is a distribution. Under equilibrium conditions the physically meaningful object is a Boltzmann-weighted ensemble over conformations,
and function frequently depends on the populations and interconversion rates of its modes rather than on the single most populated one. A predicted structure is a point estimate. It carries no populations, no relative free energies, and no exchange kinetics — none of which can be recovered from the file by inspection.
Whether current predictors can be coaxed into sampling more of that distribution is an open and actively contested question, and it is worth showing the contest rather than picking a side. Wayment-Steele, Kern and colleagues reported that clustering a multiple-sequence alignment by sequence similarity enables AlphaFold2 to sample alternative states of known metamorphic proteins with high confidence [15]. Schafer, Lee, Chakravarty and Porter, in a Matters Arising response, reported that on their test cases the same clustering procedure underperforms random sequence subsampling and therefore lacks predictive and explanatory power [16]. As of writing, a careful reader should treat “sequence clustering predicts alternative conformations” as disputed in the primary literature, not established.
Fold switching, and the silence of the confidence metric
Chakravarty and Porter tested AlphaFold2 on 98 fold-switching proteins, which adopt at least two distinct yet stable secondary and tertiary structures. They reported that 94% of predictions captured one experimentally determined conformation but not the other, and — the more consequential finding — that estimated confidences were moderate to high for 74% of fold-switching residues [17].
That second number is the one to sit with. The failure was not flagged. A confidence metric trained to predict agreement with a single deposited structure behaves exactly as designed when a protein has two: it reports high confidence in the one it found. Calibration for the question asked is not calibration for the question a biologist has.
Disorder
Wright and Dyson argued in 1999 that a substantial class of functional proteins are intrinsically unstructured under physiological conditions, and that the structure-function paradigm needed reassessment to accommodate them [19]. Ruff and Pappu, examining AlphaFold’s proteome-scale output, noted that regions of very low confidence overlap intrinsically disordered regions, that over 30% of proteome regions are disordered, and that low-confidence predictions in such regions should not be read as conformational descriptions of them [20].
The correct summary is narrow and useful: low pLDDT is an informative disorder signal and a useless ensemble description. It tells you the model does not expect a single well-defined conformation. It does not tell you what the ensemble is, how compact it is, how it responds to a binding partner, or whether it undergoes coupled folding on binding.
Mutations
Buel and Walters asked directly whether AlphaFold2 predicts the structural impact of missense mutations, and concluded that it falls short, arguing that the model is insensitive to structure-disrupting substitutions because its predictions rest on wild-type and homologous sequence signal rather than on a database of such mutations [18].
There is a structural reason for this that follows from the earlier section, and it is not a defect of engineering. The information source is the alignment. A single point substitution changes one residue in one sequence in an alignment that may contain thousands. The method is, by construction, close to blind to precisely the perturbation that matters clinically. That is a statement about predicting structural change; predictors of variant pathogenicity that consume structural features are making a different and separately evaluable claim, and should not be treated as evidence for this one.
Complexes, ligands and mechanism
Baek and colleagues reported that a three-track network could generate protein-protein complex models from sequence without a separate docking step [10], and Abramson and colleagues subsequently described AlphaFold 3, which predicts structures of complexes spanning proteins, nucleic acids, small molecules, ions and modified residues [22]. These are substantive extensions, and the appropriate posture toward them is the CASP posture: performance is a per-system empirical question, established on defined evaluation sets, not a general capability to be assumed for an arbitrary complex of interest.
Underneath all of this sits the plainest limitation. A structure is not a mechanism. Enzyme catalysis is a story about transition states, strain, protonation states, ordered water and motion between states; a ground-state coordinate file constrains hypotheses about it and does not supply one. Terwilliger and colleagues framed predictions as valuable hypotheses that accelerate but do not replace experimental structure determination, and specifically cautioned about details involving interactions not included in the prediction [14]. Agarwal and McShan’s review makes the same point as a working recommendation: integrate models with SAXS, NMR, cryo-EM and X-ray data and treat them as hypotheses rather than as de facto ground truth [21].
The loft is not the hull
It is worth being precise about what has been settled, because both overstatement and dismissal are available and both are wrong.
Settled: for single-chain proteins with adequate evolutionary information, computed models can match experimental structures closely enough that the difference approaches experimental disagreement, and this was demonstrated under a blind, independently assessed protocol rather than self-reported [8, 9, 10]. Also settled, and often underrated: the models emit calibrated per-residue confidence, which makes their output usable as evidence with a stated reliability rather than as an oracle [11, 12].
Interpretation, not fact: that this constitutes a solution to the protein folding problem. On Dill and MacCallum’s three questions, the third has largely been answered and the first has not [4]. We can now produce the structure; we still cannot read the physical code off the sequence, because the working method does not use one. What was built is an extraordinarily good interpolator over the record of sequences that already folded.
Open: whether ensembles and their populations can be predicted with calibrated uncertainty; whether any method can become sensitive to single substitutions when its signal is an alignment; how performance degrades outside the deep-alignment regime, for designed, orphan and rapidly evolving sequences; and whether confidence metrics can be made to fail loudly on the cases where they currently fail quietly, as fold switching shows they do [17].
In a mould loft the shipwright lays the hull’s lines out at full size on the floor, bends battens through the offsets, and shifts the lead ducks until every section, waterline and buttock agrees with every other. When the curves are fair, the templates come off the floor, and they are true. They are also, still, flat ply. They record what the hull’s sections must be; they do not float, they do not carry a load, and they say nothing about how the vessel behaves in a seaway. The templates are indispensable and they are not the ship. That is roughly where structure prediction now stands: the lines are faired, the templates are good, and the work of building and sailing has not been done by drawing them.