Ask a materials scientist which approach to materials discovery is “best” and the honest answer is that the question conflates three activities that solve different parts of the same problem, plus a fourth activity — degradation testing — that decides whether any of it mattered. Trial-and-error and combinatorial synthesis physically make new compositions and read off what they do. Computational high-throughput screening, most visibly the Materials Project’s density-functional-theory (DFT) database, filters an enormous space of hypothetical structures before anyone touches a crucible [1]. Machine-learning-guided autonomous laboratories, such as the University of California, Berkeley’s A-Lab, attempt to close the loop between a predicted candidate and a physically realized sample without a human deciding what to try next [3]. And degradation and qualification testing — accelerated aging, corrosion chambers, fatigue rigs — is the stage every one of the first three approaches has to pass through before a material reaches a product, because a structure’s ground-state stability says nothing about how it behaves after ten years of thermal cycling or salt exposure.
This article compares the three discovery approaches, and separately treats degradation testing as the filter none of them replaces. The comparison is made on dimensions that are genuinely measurable across methods — discovery rate, validated hit rate, and cost per candidate — rather than on a single aggregate “which one wins” score, because no such score exists and any claim that one does should be read as marketing rather than measurement.
What each approach actually does
Trial-and-error and combinatorial synthesis. The oldest approach is still the most literal: physically make a material and measure what it does. Its modern, systematized form is combinatorial materials science — depositing a gradient thin-film “library” across a single wafer so that composition varies continuously across its surface, then characterizing the whole gradient with automated probes. The approach traces to Peter Hanak’s 1970 “multiple-sample concept” and was formalized as a search strategy by Xiang and colleagues in 1995; it remains in active use today for systems where no computational model has reliable training data, particularly multinary alloys, doped oxides, and systems with strong electron correlation that DFT handles poorly [6]. Its defining property is that it does not require you to already know what to look for computationally — it explores real chemistry and real microstructure, including metastable phases and kinetic products that a ground-state-energy calculation would never surface.
Computational high-throughput screening. The Materials Project and comparable databases (AFLOW, NIST’s JARVIS-DFT) start from tens of thousands of experimentally known structures in the Inorganic Crystal Structure Database, relax each one computationally, and calculate electronic, mechanical, and thermodynamic properties at scale using DFT [1]. This is the core method underwritten by the U.S. Materials Genome Initiative, launched in 2011 with the explicit, quantified goal of cutting the time from discovery to deployment in half [2]. The approach’s strength is throughput per dollar: a DFT relaxation and property calculation on a known crystal structure costs a small fraction of what a synthesis-and-characterization cycle costs, so screening can rule out — or rank — millions of candidates before any physical sample exists. Its sharp limit is that DFT approximates the exchange-correlation functional, so its energies, band gaps, and phase-stability calls carry systematic errors that are well characterized for some chemistries and poorly characterized for others; a structure that DFT calls thermodynamically stable is a hypothesis about synthesizability, not a synthesis.
Machine-learning-guided autonomous laboratories. The newest layer trains models on the output of the first two: DeepMind’s GNoME model, trained on Materials Project data and released in 2023, used graph neural networks in an active-learning loop to propose 2.2 million candidate crystal structures, of which the model classified about 380,000 as thermodynamically stable — described by the authors as roughly 800 years’ worth of stable-structure knowledge compressed into one release [4]. Berkeley’s A-Lab paired that predicted-candidate stream with a robotic synthesis platform: over 17 days of continuous, largely unattended operation it attempted 58 synthesis targets selected from Materials Project and GNoME phase-stability data and, by its authors’ account, successfully realized 41 of them, with recipes proposed by literature-trained language models and refined through active learning grounded in thermodynamics [3]. The pitch for this category is throughput at the synthesis stage itself, not just at the screening stage: a system that runs overnight without a graduate student deciding the next experiment.
That pitch has been directly contested, which matters more here than in most technology comparisons because the contested number is the discovery-rate figure the category is sold on. A 2024 perspective by Anthony Cheetham and Ram Seshadri re-examined a sample of GNoME’s “stable” output and argued that many of the proposed structures are better described as trivial dopant variants, symmetry-broken duplicates of already-known phases, or chemically implausible high-element-count compositions rather than new, synthesizable materials [5]. Subsequent reporting on crystallography databases more broadly found that a nontrivial fraction of entries across multiple high-throughput datasets are near-duplicate structures rather than independent discoveries [9]. This is not a claim that GNoME or A-Lab produced nothing — A-Lab’s own authors later clarified that their “novel” framing meant new to the prediction platform’s training set, not necessarily new to science, after outside groups raised questions about phase identification from diffraction data. It is a claim that the headline discovery-rate number for the autonomous-lab category is disputed by domain specialists, and any comparison that repeats it uncritically is repeating a contested figure as settled fact.
Discovery rate: comparable only within, not across, categories
“Discovery rate” is not one number. A combinatorial synthesis campaign might characterize a few
hundred distinct compositions on one wafer in a single run, each with a real, physically measured
property value. A DFT screening pass processes a database on the order of
The rate comparison that is defensible is narrower: given a fixed synthesis and characterization budget, does DFT or ML pre-screening improve the fraction of attempted syntheses that succeed, relative to unscreened trial-and-error? A-Lab’s reported 41-of-58 success rate over 17 days, drawing its candidate list from Materials-Project- and GNoME-informed phase-stability data, is presented by its authors as evidence that pre-screening does raise the hit rate for a robotic synthesis campaign relative to naive attempts [3]. That is a real, checkable claim about one system under one protocol. It is not evidence that GNoME’s underlying stable-structure count of 380,000 is itself mostly correct, and the Cheetham–Seshadri critique is specifically about that separate number [5].
Validated hit rate is the number that actually separates the approaches
If discovery-rate comparisons collapse under unit mismatch, validated hit rate — the fraction of candidates from a given method that survive independent experimental confirmation — is closer to an apples-to-apples measure, though it still requires stating what “validated” means for each method.
For combinatorial synthesis, validation is close to immediate: the material was physically made, and characterization (X-ray diffraction, electron microscopy, electrical or magnetic measurement) confirms what phase and composition actually formed, subject to the usual measurement uncertainty of those techniques. The category’s weakness is coverage, not confidence — it validates a few hundred to a few thousand compositions per campaign, not millions.
For DFT screening, “validated” means an independent experimental synthesis matches the predicted structure and its predicted stability ranking, which happens for only a subset of the candidates that ever get pursued, because most database entries are never attempted experimentally at all. The honest DFT hit-rate statement is therefore conditional: among predicted-stable candidates that someone actually tried to synthesize, what fraction were realized as predicted? Neither this article nor, as far as searches for this piece could establish, any single published source states that number across the whole Materials Project catalog; it can only be quoted per targeted campaign, such as A-Lab’s own 41-of-58 [3].
For ML-guided autonomous labs, validated hit rate is the number under direct dispute. A-Lab reports roughly 70 percent of its 58 attempted targets synthesized as intended over its reported run [3]. GNoME’s own 380,000-structure “stable” count is a computational classification, not an experimental validation count, and the critique from Cheetham and Seshadri argues that an unknown but material fraction of that computational count would not survive independent scrutiny for novelty or synthesizability even before anyone attempts to make them [5]. Reporting on duplicate structures across high-throughput databases generally reinforces that caution without pinning down a single cross-database duplicate rate [9].
Writing the ratio out this way is useful precisely because it exposes the trap: the denominator
means something different for each method. For combinatorial synthesis, the denominator and
numerator are nearly the same population, because everything proposed gets made. For DFT and ML
screening, the denominator is enormous and the numerator is whatever small subset anyone bothered to
attempt — so a favorable-looking
Cost per candidate: where the categories genuinely diverge
Cost is the dimension where the three approaches are least ambiguous, because each has a real, quotable unit cost, even though the units differ.
A single DFT relaxation-and-property calculation on a known or near-known crystal structure typically consumes compute measured in core-hours to a modest number of node-hours — cheap enough that a national-lab-scale project can process tens of thousands of structures as a matter of course [1]. This is the strongest, least contested comparative claim in the field: DFT screening is orders of magnitude cheaper per candidate than physically synthesizing and characterizing that same candidate, full stop, independent of any dispute about how many of the resulting predictions are actually good.
A combinatorial synthesis run has a materially higher per-candidate cost — instrument time on a sputtering or pulsed-laser-deposition chamber, characterization-probe time across the gradient, technician oversight — but that cost buys physically confirmed compositions rather than predictions, and it is the only one of the three methods that can discover behavior the underlying computational models were never trained to represent, because it does not depend on a model at all [6].
An autonomous lab’s marginal cost per attempted synthesis sits closer to conventional trial-and-error than to DFT screening — A-Lab still consumes furnace time, reagent stock, and robotic-arm cycles for every attempt — but its claimed advantage is on the other side of the ledger: fewer human-decision cycles between attempts, and continuous overnight operation without a researcher choosing the next experiment by hand [3]. Whether that translates into lower cost per validated material, rather than merely lower cost per attempted synthesis, is exactly the question the Cheetham–Seshadri critique leaves open, because a synthesis attempt guided by a disputed prediction is not obviously cheaper, per confirmed discovery, than one guided by a combinatorial gradient or a conservative DFT ranking [5].
The overclaim to watch for, stated plainly
The most common overclaim in vendor and press coverage of this space is a chained inference: (1) GNoME predicted 380,000 stable structures, therefore (2) materials science now has 380,000 new materials, therefore (3) AI has out-produced a century of human materials discovery in one release. Step 1 is a real, checkable computational output [4]. Step 2 substitutes “predicted stable by one model” for “discovered,” which the Cheetham–Seshadri critique and subsequent duplicate-structure reporting directly contest for a material fraction of that count [5] [9]. Step 3 compounds an already-contested count with a historical comparison that has no agreed unit of measurement on either side. None of this means GNoME’s underlying method is worthless — active-learning graph networks trained on Materials Project data are a real methodological contribution, and A-Lab’s synthesis results, even read conservatively, show pre-screening can raise a robotic synthesis campaign’s hit rate above naive attempts [3]. It means the size of the claimed discovery, specifically, is the part under dispute, and a comparison piece that repeats 380,000 as a count of new materials without that caveat is repeating an industry press-release framing rather than a verified result.
Where all three approaches hand off to the same filter: degradation
Discovery, however it is done, produces a candidate structure or composition. None of the three approaches compared above tells you whether that material survives its intended service life, and this is not a gap that faster discovery closes — it is a categorically separate question answered by accelerated-aging and degradation testing.
Two worked examples make the handoff concrete. In lithium-ion batteries, the dominant capacity-fade mechanisms — solid-electrolyte-interphase (SEI) growth on the anode, lithium plating under fast charging, particle cracking from repeated volume change, and transition-metal dissolution from the cathode — are properties of a cell’s long-term cycling behavior, not of a candidate electrode material’s ground-state DFT energy or its GNoME stability score [8]. A material can be correctly predicted as thermodynamically stable and still fail in service within a few hundred cycles because of an interfacial degradation pathway that only shows up under repeated electrochemical cycling. In corrosion testing, ASTM B117 — first published in 1939 and still the most widely used salt-fog corrosion protocol worldwide alongside ISO 9227 — exposes coated or uncoated metal coupons to a continuous 5-percent-sodium-chloride fog at 35 degrees Celsius for a specified duration to produce a relative corrosion-resistance ranking [7]. The standard’s own documentation is explicit that its results are relative rankings within one test chamber and have seldom correlated reliably with performance in real, uncontrolled environments — a caveat that matters for any claim that a laboratory corrosion result predicts field service life.
The mechanical-metamaterials case sharpens the same point from a different angle. A computationally designed lattice — an octet truss or similar architected structure — can be optimized in silico for a target stiffness-to-weight ratio, and DFT-adjacent atomistic methods can predict the base material’s elastic constants reasonably well. Whether the as-manufactured lattice reaches that target depends on print-defect statistics, strut-level stress concentration, and fatigue behavior under cyclic loading that only a physical tensile or fatigue rig captures, because those failure modes emerge from manufacturing variance and load history, not from the unit cell’s idealized geometry.
This Palmgren–Miner-style cumulative-damage sum, where
What a fair comparison actually says
Trial-and-error and combinatorial synthesis remain the only approach that generates genuinely new experimental data outside a model’s training distribution, at a real but bounded per-candidate cost, and it is where the ground truth used to train every computational model ultimately comes from. Computational high-throughput screening is unambiguously cheaper and faster per candidate at the screening stage, with a real, well-documented history running from the Materials Genome Initiative’s 2011 founding through the Materials Project’s database [2] [1], but it inherits DFT’s known systematic errors and produces hypotheses, not confirmed materials. ML-guided autonomous laboratories compress the loop between a computational candidate and a physical attempt and have shown, in at least one well-documented case, a higher per-attempt success rate than unguided synthesis [3] — but the categorical claim that this approach has multiplied the rate of genuine materials discovery by orders of magnitude is specifically the part that domain specialists have publicly and substantively disputed [5] [9]. And every one of the three, regardless of how a candidate was found, still has to clear the same degradation-testing gate — battery cycling, salt-fog corrosion exposure, fatigue loading — before anyone can say the material does what it was chosen to do over years rather than in one measurement.
Prediction, stated as a prediction, not a fact: over the next five years (through 2031), expect autonomous-synthesis platforms to become common infrastructure at national labs and large materials companies specifically for narrow, well-characterized chemical families — battery cathodes, catalytic alloys — where training data is abundant, rather than as general-purpose discovery engines across all of materials space. This assumes continued public funding at roughly Materials-Genome-era levels and no major redirection of DFT-database maintenance funding. An observable indicator would be a rising count of peer-reviewed papers reporting autonomous-lab-discovered materials that pass independent third-party synthesis replication, not just internal validation by the group that built the model. The prediction would be disconfirmed if, by 2031, the fraction of ML-predicted “stable” structures confirmed by independent groups remains at or below the level implied by the Cheetham–Seshadri critique’s 2024 assessment, with no measurable improvement in independently replicated hit rate despite continued scale-up of prediction volume [5].