Four forks, not one method

“AI for science” reads like a single project with a single scoreboard. It is not. Look closely at any one of the domains where machine learning has produced a genuinely useful scientific result, and the same pattern repeats: two or three methodologically distinct approaches are running at once, solving related but not identical problems, validated against different baselines by different standards of proof. Weather forecasting has physics-based numerical prediction, neural-network forecasting, and hybrid systems that couple the two. Materials and drug discovery have generative candidate search, exhaustive high-throughput screening, and active-learning-guided robotic experimentation. Protein science has end-to-end structure prediction and physics-based molecular dynamics simulation, which answer different questions about the same molecule. And a newer contest sits above all three: general-purpose “AI scientist” agents against the narrow, task-specific models that produced most of the field’s credible results so far.

This article works through each fork in turn. The organizing discipline is the one this publication holds to everywhere: separate what has actually been demonstrated, under what baseline, assessed by whom, from what is a vendor’s characterization of its own system, from what is this article’s analysis of the gap between the two. Where two approaches have been tested under matched conditions — same task, same withheld data, same independent assessor — say so plainly and report the result. Where they have not, because the benchmarks, the domains, or the goals genuinely differ, say that too, rather than manufacturing a ranking the evidence does not support. That second case turns out to be the more common one.

Weather: a rare case of clean, replicated comparison

Weather forecasting is the domain where a head-to-head comparison between physics-based and data-driven methods is actually possible, and it has now been run by more than one independent group against the same kind of baseline. That makes it worth starting here — not because it is representative of the rest of the field, but precisely because it is not, and the contrast is instructive.

ADVERTISEMENT

Operational numerical weather prediction solves a form of the primitive equations of fluid dynamics on a global grid, using observed atmospheric conditions as the initial state. The European Centre for Medium-Range Weather Forecasts’ deterministic system, HRES, and its probabilistic ensemble system, ENS, are the standard incumbents against which new methods are typically checked, because they are public, operational, and have been continuously verified for decades. Three independent groups have now trained neural networks to forecast directly from the same reanalysis data used to initialize these physical models, and evaluated the results against ECMWF’s own operational output.

DeepMind’s GraphCast, a graph neural network trained on ERA5 reanalysis data, was reported to significantly outperform the most accurate operational deterministic system on 90% of 1,380 verification targets, while predicting hundreds of weather variables ten days out, globally, in under one minute [1]. Huawei’s Pangu-Weather, a three-dimensional neural network trained similarly, was described by its authors as the first AI-based method to outperform traditional numerical weather prediction, reporting better latitude-weighted RMSE and anomaly correlation across all measured variables and all lead times from one hour to one week [2]. DeepMind’s later GenCast, a diffusion-based ensemble model, reported greater skill than ECMWF’s ENS ensemble on 97.4% of 1,320 evaluated targets, while also better predicting extreme weather, tropical cyclones, and wind-power-relevant conditions [3].

Read carefully, these are three separate results, not one number repeated three times, and they should not be stacked into a single “AI beats physics” scoreboard. GraphCast and Pangu-Weather are deterministic point forecasts checked against HRES using RMSE and anomaly correlation; GenCast is a probabilistic ensemble checked against ENS using different skill measures suited to ensembles. A deterministic forecast and a probabilistic ensemble forecast are not interchangeable, and “beats the baseline on 90% of targets” in one paper is not directly comparable to “beats the baseline on 97.4% of targets” in another, because the target sets, metrics, and task itself differ. What is genuinely established is narrower and more solid: for medium-range forecasting at fixed lead times, initialized from the same observational analysis, multiple independently built neural networks now match or exceed a specific, named, operational physics-based baseline on the majority of quantities each study measured. That is a real result, replicated across three organizations using three different architectures — unusually strong evidence by this field’s standards.

It is also a narrower result than “neural networks have replaced physics in weather science,” and Google’s own follow-up work makes the boundary explicit. NeuralGCM couples a differentiable, physics-based dynamical core to learned components that stand in for processes too fine-grained to resolve directly, such as cloud formation [4]. It is not a competitor to GraphCast or GenCast in the sense that GraphCast competes with HRES; it is a third architecture that keeps the physical solver rather than discarding it. Its reported advantage is not raw medium-range skill, where it matches rather than dramatically exceeds pure neural approaches, but stability over far longer horizons: accurately tracking climate metrics such as global mean temperature over multiple simulated decades, and reproducing emergent phenomena such as realistic tropical cyclone behavior, at a fraction of the computational cost of a conventional circulation model. Pure neural weather models are trained and evaluated on medium-range forecasts of days to a couple of weeks; nothing in the GraphCast, Pangu-Weather, or GenCast results claims stability over the decades-long horizons climate projection requires, and that is a different validation problem with no equivalent blind leaderboard yet in place. “AI beats physics for medium-range weather forecasting” has real, replicated, independently checked evidence behind it. “AI beats physics for climate simulation” is not a claim any of these papers make, and the honest comparison is that hybrid architectures currently occupy that longer-horizon niche precisely because pure data-driven models have not been shown to belong there.

A physics-based weather-simulation console and a neural-network forecasting rack on adjoining benches, a bridging cable caught mid-plug between the two systems
Figure 1. GraphCast, Pangu-Weather and GenCast were each verified against the same operational baseline ECMWF runs, one of the few genuinely clean head-to-head comparisons in AI for science; NeuralGCM instead couples a physics solver to learned components rather than replacing it.

Materials and drug discovery: three paradigms that solve different problems

Move from weather to materials and drug discovery, and the clean baseline disappears. Three approaches are active here, and they are frequently discussed as though they compete for the same prize. They do not: they operate on different search spaces, pursue different objects, and are checked against different criteria for success.

ADVERTISEMENT

Generative candidate search proposes new compositions that were never previously enumerated. DeepMind’s GNoME used large-scale graph-network training to predict the thermodynamic stability of candidate inorganic crystal structures, and applied at scale it proposed 2.2 million new crystal structures, of which roughly 380,000 were predicted to be stable enough to be promising synthesis candidates [5]. That is worth setting against what existed before it: prior computational screening, accumulated over years, had reached roughly 48,000 stable predicted materials, against a base of around 20,000 materials experimentally characterized by traditional means over preceding decades. On the community benchmark MatBench Discovery, GNoME-style training reportedly improved discovery-rate precision from around 50% to over 80%. But a predicted stable structure is a computational claim, not a synthesized material, and the disclosing lab’s own account of independent confirmation is modest relative to the scale of the prediction: external laboratories have independently created 736 of the newly predicted structures in concurrent work — a genuine validation signal, but a small fraction of the 380,000 structures flagged as most stable. The honest description is that GNoME expanded the computational search space by roughly an order of magnitude and that a few hundred of its predictions have since been confirmed by independent synthesis, not that hundreds of thousands of new materials now exist.

High-throughput virtual screening takes the opposite approach: rather than proposing new compositions, it exhaustively scores a library that already exists. The clearest demonstration is ultra-large library docking, in which researchers computationally docked a library of 170 million make-on-demand compounds, built from 130 well-characterized reactions, against two protein targets: AmpC beta-lactamase and the D4 dopamine receptor [8]. Against AmpC, 99 million library members were scored, 51 top-ranked compounds were selected and 44 synthesized, yielding an 11% hit rate and, after optimization, a 77-nanomolar inhibitor among the most potent non-covalent AmpC inhibitors then known. Against D4, 138 million compounds were scored, 589 selected and 549 synthesized, and 122 bound the receptor — a hit rate that fell roughly monotonically as docking score worsened. That monotonic relationship is itself the interesting result, because it let the authors extrapolate a library-wide estimate rather than test everything directly:

N^hits  ≈  ∑sh(s) N(s) \hat{N}_{\text{hits}} \;\approx\; \sum_{s} h(s)\, N(s) ↗

where h(s)h(s)↗ is the hit rate measured directly among compounds actually synthesized and tested at docking-score bin ss↗, and N(s)N(s)↗ is the number of untested library members scored into that bin. Applying this to the D4 results produced an estimate of roughly 453,000 ligands across the full 138-million-compound library — a projection under an assumed score-to-hit-rate relationship, not a directly counted total. Confirmed hits included a 180-picomolar, subtype-selective D4 agonist, one of 81 new chemotypes with no prior literature precedent, 30 of which showed submicromolar activity. This is screening logic, not generative logic: nothing here proposes a molecule outside the pre-enumerated library, and the achievement is exhaustive search of a known space rather than invention of a new one.

A generative candidate-search terminal ejecting a crystal-structure readout card beside a high-throughput robotic screening arm mid-dispense over a well plate
Figure 2. GNoME proposes candidate crystal structures a physics solver never tested by synthesis; ultra-large library docking instead scores a library that already exists. They answer different questions about different kinds of molecule.

Active-learning-guided experimentation is the third paradigm, and it is the only one of the three that closes a physical loop: propose, synthesize, characterize, and revise, with a robotic laboratory in the loop rather than a one-shot computational output. The clearest public demonstration is the A-Lab at Lawrence Berkeley National Laboratory, which combined phase-stability predictions drawn from the Materials Project and GNoME, synthesis-recipe proposals generated by literature-trained language models, and an active-learning loop grounded in thermodynamics to guide a robotic solid-state synthesis and X-ray diffraction characterization system [6]. Over 17 days of continuous, largely unattended operation, running 353 experiments, the A-Lab successfully synthesized 36 of 57 targeted novel inorganic compounds spanning 33 elements and 40 structural prototypes — a 63% success rate, at a pace exceeding two new materials per day.

A robotic arm carrying a ceramic crucible toward an open furnace mouth, an XRD readout station standing just behind it, the crucible not yet inside the furnace
Figure 3. Active-learning-guided synthesis closes a loop that generative search and library screening do not: propose, synthesize, characterize, and revise. That loop is also where the field's sharpest reproducibility dispute so far has landed.

That result also produced this domain’s sharpest reproducibility dispute so far, and it belongs in a fair comparison precisely because it does not flatter the paradigm that looks strongest on paper. A subsequent independent re-analysis by Robert Palgrave at University College London and Leslie Schoop at Princeton, reported by Chemistry World, challenged a substantial share of the claimed novel compounds, arguing that roughly two-thirds of the reported new structures were, on closer inspection, already-known disordered compounds that had been misidentified as new ordered phases because of inadequate diffraction-pattern refinement, which the critique’s authors characterized bluntly as novice-level [7]. In the critique’s strongest form, Palgrave concluded it was likely that no genuinely new compounds had been discovered at all. The original team’s response, reported in the same account, acknowledged that a human expert could perform a higher-quality refinement than the automated pipeline had, while maintaining that the paper’s central claim was about demonstrating autonomous-laboratory capability rather than guaranteeing every reported compound was novel. Readers should treat this as an open, contested disagreement between named experts rather than a settled correction in either direction — but it is a fact about the state of the evidence that belongs in any comparison of these three paradigms, and it is a useful caution against treating the most impressive-sounding demonstration as automatically the most reliable one.

Put the three side by side and the reason a ranking is not available becomes structural rather than a gap in the evidence. GNoME proposes inorganic crystal compositions and is checked against thermodynamic stability calculations and a small, growing count of independent syntheses. Ultra-large library docking scores drug-like small molecules from an already-enumerated library and is checked against measured binding affinity in synthesized compounds. The A-Lab attempts to close the loop from prediction to physical realization for inorganic powders and is checked against diffraction-confirmed synthesis, a claim now under direct dispute for a meaningful share of its results. These are different molecules, different objects of proof, and different failure modes — a ranking of “which AI approach works best for discovery” would require collapsing distinctions that are the actual substance of the comparison.

ADVERTISEMENT

Protein structure: what prediction settles, and what it cannot address

Protein science offers the cleanest single result in this entire survey, and also the clearest illustration of why a clean result in one narrow sense can be quietly misread as settling a broader question it never addressed.

AlphaFold2 was assessed in CASP14, the blind, independently organized structure-prediction assessment in which competing methods predict structures for proteins whose experimentally solved coordinates are not yet public, and are scored by assessors with no stake in any team’s outcome. Against that standard, AlphaFold2’s median backbone accuracy was 0.96 angstroms of Cα root-mean-square deviation from the experimental structure, against 2.8 angstroms for the next-best method, and its median all-atom accuracy was 1.5 angstroms against 3.5 angstroms for the next-best method [9]. What that root-mean-square-deviation figure actually measures is worth stating explicitly, because the number is often quoted without the definition it depends on:

RMSD=1N∑i=1N∥xi−x^i∥2 \text{RMSD} = \sqrt{\frac{1}{N} \sum_{i=1}^{N} \lVert x_i - \hat{x}_i \rVert^2} ↗

where xix_i↗ is the coordinate of atom ii↗ in the experimentally solved structure, x^i\hat{x}_i↗ is the corresponding coordinate in the predicted structure after optimal superposition of the two, and NN↗ is the number of atoms compared. This is a statement about the average positional deviation of one predicted, static conformation from one experimentally solved, static conformation. It is not, by construction, a statement about how the molecule moves, how quickly it interconverts between conformations, or what its folding pathway looks like — and that distinction is the entire content of what comes next.

A molecular-dynamics compute cluster with active cooling manifolds running beside a structure-prediction inference workstation ejecting a single static-model card
Figure 4. A structure predictor outputs one static conformation and stops; a molecular-dynamics cluster keeps running, computing how that same molecule moves. They are not competing on the same measurement.

Molecular dynamics simulation answers a different question about the same molecule: given a physics-based force field describing interatomic interactions, how does the structure evolve over time. Reviews of the ML-augmented state of the field describe the tools built for exactly this purpose — neural networks trained to predict quantum-mechanical energies and forces more cheaply than direct calculation, coarse-grained simulation methods, and techniques for extracting free-energy surfaces and kinetic rates from simulated trajectories [11]. None of these tools compete with AlphaFold2 on CASP’s benchmark, because CASP does not ask for a trajectory, a free energy, or a rate; it asks for one static structure, scored against one static experimental answer.

Where the two have been made to overlap, the overlap is narrow and explicitly provisional. A widely cited demonstration showed that subsampling the multiple sequence alignment fed to AlphaFold2 — deliberately reducing the evolutionary information the model sees — could coax it into generating structures spanning multiple conformational states for some transporters and receptors, with models at the extremes of the resulting distribution reaching an average template-modeling score of 0.94 against experimentally known alternate conformations [10]. That correlation with real conformational flexibility is a genuinely interesting emergent property. It is also, in the authors’ own words, “a methodological hack, rather than an explicit objective,” and it works best for proteins whose alternate conformations were not already in AlphaFold2’s training data — proteins the model has effectively memorized behave far less flexibly under the same procedure. This is a narrow, acknowledged workaround, not a general substitute for physics-based simulation of molecular motion. Structure prediction and dynamics simulation remain two tools answering two different questions, one validated by a blind static-structure benchmark, the other by agreement with physical theory and, where available, measured kinetics.

General-purpose reasoning agents against narrow specialists

Every result described so far — GraphCast against HRES, GNoME against synthesized crystals, AlphaFold2 against CASP14 — is a narrow, single-purpose model, built for one task and checked against one benchmark run by people with no stake in the outcome. The newest entrant in AI for science is built differently: general-purpose systems that use a large language model as a reasoning core, wrapped in tools for literature search, code execution, and laboratory control, aimed at the open-ended work of proposing and pursuing a research question rather than solving one fixed task.

Coscientist, built at Carnegie Mellon University on GPT-4, demonstrated autonomous experimental design and execution across six tasks: designing synthesis procedures for common pharmaceuticals including aspirin, acetaminophen, and ibuprofen; controlling robotic liquid handlers to produce visual patterns in laboratory well plates; identifying unknown liquids via spectrophotometer readings; writing code to operate laboratory hardware; and, most notably, planning and executing real Suzuki and Sonogashira palladium-catalyzed cross-coupling reactions, the chemistry that won the 2010 Nobel Prize [12]. When a coding error caused a heating device to malfunction, the system diagnosed the fault, consulted the device’s technical manual, corrected its own code, and retried without human intervention. On one task, identifying unknown colored liquids, researchers noted it required a nudge in the right direction. Six demonstrated tasks, largely successful, is a genuinely impressive account of tool-using capability. It is not a blind benchmark result in the CASP or WeatherBench sense; it is a set of case studies reported by the team that built the system.

Google’s AI co-scientist, built on Gemini, extends the same idea with a more elaborate architecture: specialized Generation, Reflection, Ranking, Evolution, Proximity, and Meta-review agents coordinated by a supervisor, using iterative self-play scientific debate and tournament-style ranking to refine hypotheses, with reported quality improving as more test-time compute is spent on the cycle [13]. Its developers describe it as designed to surpass unassisted human experts on complex problems — a claim from the system’s own developers, to be read as a vendor characterization rather than an independently adjudicated result. Its most concrete evidence is a small set of biomedical case studies. In one, proposed drug-repurposing candidates for acute myeloid leukemia were subsequently confirmed in laboratory testing to inhibit tumor viability at clinically relevant concentrations. In another, the system generated fifteen hypotheses about liver fibrosis, identifying three novel epigenetic targets, two of which showed anti-fibrotic activity without toxicity in human hepatic organoids. In a third, on antimicrobial resistance, the system proposed investigating an interaction between phage-inducible chromosomal islands and phage tails that matched a mechanism researchers had already discovered — but had already experimentally validated before revealing the problem to the system, making this a demonstration of reconstructing a known, unpublished result rather than a prospective discovery. The developers’ own account notes a small sample size in expert assessments, no capacity to design a full clinical trial, no modeling of drug bioavailability or pharmacokinetics, and inherited limitations from the underlying language model.

A general-purpose reasoning agent's desk with a robotic arm reaching between several different instruments, beside a narrow specialized model's bench where a single instrument carries a validation tag being affixed
Figure 5. A narrow specialized model is validated against one blind, independently run benchmark. A general-purpose scientific agent is currently judged by a handful of demonstration case studies reported by the lab that built it — a different, and thinner, standard of proof.

The honest comparison between general-purpose agents and narrow specialists is not about which produces better science; on present evidence it is about which is validated by what standard, and those standards are not currently commensurable. AlphaFold2’s claim rests on CASP14, decades old, independently organized, genuinely blind. GNoME’s claim rests partly on the community-run MatBench Discovery benchmark and partly on a small, publicly counted number of independent syntheses. Even the deployment side of narrow models — the part closest to clinical use — is less uniformly rigorous than the phrase “narrow, validated model” implies. A 2025 analysis of the 168 machine-learning-enabled medical devices the FDA authorized in 2024 found that 94.6% were cleared through the 510(k) pathway, which requires demonstrating substantial equivalence to an already-cleared predicate device rather than a prospective clinical trial, and that only 29.2% publicly reported both sensitivity and specificity, with just 15.5% disclosing race or ethnicity data about the populations tested [14]. Regulatory authorization is a real form of validation, but it is not automatically equivalent to a blind accuracy benchmark, and the gap between “FDA-cleared” and “independently shown to work at a stated accuracy on a representative population” is an open transparency problem for narrow medical AI too, not only for speculative general-purpose agents. Against that backdrop, the agents’ demonstration case studies are not a lesser form of the same evidence CASP or the FDA’s device database provide; they are a different, thinner category of evidence, self-reported by the labs that built the systems, with at most a handful of cases and at least one confirming a result already known rather than a prospective discovery. That is not a reason to dismiss what these agents have shown. It is a reason not to rank them against narrow specialized models as though both were measured by the same yardstick, because they are not.

The validation standard is the real variable

Step back across all four forks and one pattern accounts for most of the apparent disagreement about which AI-for-science approach is winning: the strength of any comparison is bounded by the independence and blindness of its assessment, and the domains surveyed here sit at different points on that spectrum. Weather forecasting has the strongest case, because three separate organizations checked their models against the same kind of long-standing, publicly operated physical baseline, on data neither model had trained on for the forecast being verified. Protein structure prediction has an equally strong case for the narrow claim CASP actually tests, precisely because CASP was built from the outset as a blind, adversarial, independently run assessment with no financial stake in any participant’s result. Materials discovery is more mixed: GNoME’s stability predictions are checked by a mostly computational, partly independent benchmark, while the A-Lab’s synthesis claims are the subject of open, unresolved expert disagreement about how many reported successes are real. Scientific-reasoning agents sit at the weakest point on the spectrum, validated mainly by demonstration case studies the building organization selects and reports itself, with one headline result turning out, on close reading, to confirm something already known rather than new.

None of this means the weaker-validated categories are lesser achievements, or that one materials dispute discredits active-learning-guided experimentation as a paradigm. It means a reader should ask, before accepting any comparison: who ran the test, did they have a stake in the outcome, and was the withheld data genuinely unseen by the system being scored. Where the answer is CASP or a matched ECMWF baseline, the comparison is close to as strong as this field currently produces. Where the answer is “the lab that built the system, in a small number of self-selected examples,” the result may still be true, but it has not yet been tested the way the field’s strongest results have.

Predictions, with the observations that would falsify them

These are forecasts, separated clearly from the sourced analysis above. Horizon: five years from today, mid-2031.

One. Independent, blind, third-party benchmarks for general-purpose scientific-reasoning agents — organized the way CASP is organized for structure prediction, with assessors who have no stake in any participating system — will emerge for at least one scientific domain. Disconfirmed if by 2031 the leading evidence for such agents’ capabilities still consists chiefly of self-reported demonstration case studies from the organizations that built them.

Two. Hybrid architectures that couple a physics-based solver to learned components, following the NeuralGCM pattern, will become the standard approach for any AI-for-science task with a long simulated horizon or a strong requirement for physical consistency, while purely data-driven models remain dominant for short-horizon, densely observed prediction tasks resembling medium-range weather forecasting. Disconfirmed if a purely data-driven model is shown to match a physics-coupled hybrid’s stability over multi-decade simulated horizons without incorporating any explicit physical constraint.

Three. At least one further high-profile active-learning-guided materials or drug-discovery result will face a public, named reproducibility challenge of the kind directed at the A-Lab, as the pace of autonomous-laboratory publication continues to outrun the pace of independent replication. Disconfirmed if the materials-discovery field establishes a routine, pre-registered independent-replication step for autonomous-lab claims before this prediction’s horizon, and no further public dispute of this kind occurs.

Four. Regulatory transparency requirements for AI-enabled medical devices will tighten to require standardized reporting of accuracy metrics and tested-population demographics, narrowing the gap documented in the FDA’s 2024 device data. Disconfirmed if the share of newly authorized devices publicly reporting both sensitivity and specificity has not measurably increased from the 2024 baseline.

What to take away

There is no single winner among the methods surveyed here, and that is not a hedge; it is the finding. Physics-informed and data-driven weather models have been cleanly compared and the data-driven models currently win the specific, narrow, medium-range comparison that has actually been run — while hybrid architectures occupy a longer-horizon niche neither pure approach has been shown to fill. Generative candidate search, high-throughput screening, and active-learning-guided experimentation solve different problems for different kinds of molecule, checked against different kinds of proof, with the most operationally ambitious of the three currently carrying the field’s sharpest open reproducibility dispute. Structure prediction and molecular dynamics simulation answer different questions about the same protein, one validated by a blind static-structure benchmark and the other by agreement with physical theory, and attempts to merge them remain narrow, acknowledged workarounds rather than a general solution. And general-purpose scientific-reasoning agents are, for now, validated by a standard of evidence categorically thinner than the blind benchmarks and regulatory review processes that took decades to build around narrower models — a gap in maturity, not necessarily a gap in eventual capability, but a gap a careful reader should not paper over. Ask, for any claim in this space, what exactly was compared against what, by whom, and under what conditions the comparison would have gone the other way. Where that question has a solid answer, treat the result as real. Where it does not, the honest response is to say so, not to rank anyway.