The dispute comes before the discovery

In November 2023, Google DeepMind announced that a graph-neural-network system called GNoME had generated 2.2 million candidate inorganic crystal structures, of which roughly 380,000 sat below the current convex hull — the standard computational proxy for thermodynamic stability — and that the underlying model had already improved the historical hit rate for stable-structure prediction from about 50% to 80% through active learning [3]. The paper itself, published the same day in Nature, called this “an order-of-magnitude expansion in stable materials known to humanity,” reported that 736 of the stable structures had already been independently realized experimentally elsewhere, and specifically flagged layered compounds and solid-electrolyte candidates as the kind of application the dataset was meant to unlock [1]. A companion paper from Lawrence Berkeley National Laboratory, timed to the same release, reported that an autonomous robotic platform called the A-Lab had taken 58 of the newly proposed compositions, worked out synthesis routes for them without human intervention, and realized 36 of 57 targets over 17 days of continuous operation [9]. Read together, the two papers described a closed loop: a model proposes, a robot attempts, and the success rate becomes the next round of training data. The press coverage at the time, and DeepMind’s own framing, treated the loop as already working.

It is not obviously working, and the argument that it is not comes from inside the field rather than from outsiders skeptical of AI. In April 2024, the materials chemists Anthony Cheetham and Ram Seshadri published a peer-reviewed perspective examining the GNoME dataset directly. Their method was blunt: they pulled ten random entries from the “Stable Structure” release and checked each one against the Inorganic Crystal Structure Database. All ten were already there, usually filed under a higher-symmetry space group than GNoME had assigned [2]. Scaling the check up, they report that non-centrosymmetric space groups make up roughly 34% of GNoME’s predictions against about 1% of the real ICSD, while the two most common space groups in the actual database — Pnma and P2₁/c, together covering about 16% of known inorganic structures — account for only about 2.26% of GNoME’s predictions [2]. Their explanation is mechanical rather than accusatory: density-functional theory at zero kelvin forces cations onto distinct crystallographic sites even when those cations are chemically similar enough, and the temperature high enough, that real crystals let them disorder onto shared sites — a terbium ion and a samarium ion of nearly identical radius, for instance, predicted as neatly separated when a real sample would mix them and collapse into a simpler, higher-symmetry structure. The dataset also contains 18,138 predicted compounds built from promethium, actinium, or protactinium and a further 23,529 built from technetium, neptunium, or plutonium — elements with negligible natural abundance and no plausible industrial use, included because the model was never told to care [2]. Cheetham and Seshadri’s summary, after examining “hundreds” of entries by their own account, is direct: “we have yet to find any strikingly novel compounds in the GNoME and Stable Structure listings” [2]. Their proposed standard for what should count as a genuine materials discovery — novelty, credibility, and utility, which they call the trifecta — is one most of the dataset fails on at least one axis, in their reading.

An XRD sample changer's carousel arm caught mid-swap between a just-measured specimen disc and the next queued one, the diffraction beam shutter only half-retracted
Figure 1. The instrument that actually settles whether a predicted structure is new: a diffraction pattern, read against everything already catalogued.Image prompt and art direction by Brecht Corbeel; generation pending.

The A-Lab synthesis claim was checked separately, and came out worse. Josh Leeman and six co-authors, in a 2024 paper devoted specifically to that question, went through all of the paper’s claimed synthetic products — their abstract puts the number at 43 — and identified four recurring shortfalls in how the diffraction data behind each claim had been interpreted [11]. The largest single problem was the same one Cheetham and Seshadri had flagged in the predictions: the A-Lab’s target compounds were computed with every element locked onto its own distinct crystallographic site, but real powders let chemically similar elements share sites, producing a higher-symmetry structure that is often already a known alloy or solid solution rather than the ordered compound the model had proposed. Leeman and colleagues’ conclusion is stated without hedging: “we find that two thirds of the claimed successful materials in Szymanski et al. are likely to be known compositionally disordered versions of the predicted ordered compounds,” and they write that the combined errors “unfortunately lead to the conclusion that no new materials have been discovered in that work” [11]. Their second major finding was a tooling gap rather than a chemistry one: automated Rietveld refinement — the standard method for matching a measured diffraction pattern against a candidate structure — is not yet reliable enough to run unsupervised at the scale the A-Lab needs, which means the very software layer meant to remove humans from the loop is the layer both critiques identify as the weak link [11].

ADVERTISEMENT

Note what the two critiques agree and disagree on before going further, because the disagreement is where the real content is. Neither group disputes that GNoME’s underlying graph-network approach works as advertised at the narrow task it was trained on: ranking candidate structures by predicted formation energy against a convex hull built from Materials Project data. Cheetham and Seshadri say so explicitly — “the underlying approach is sound” [2]. What both groups dispute is the translation from that narrow, low-temperature computational result into the public claim of 2.2 million new materials, and specifically the step where “thermodynamically favorable at zero kelvin, with every atom on its own site” gets reported as “novel,” when a huge fraction of those structures are, empirically, existing compounds that a disorder-blind calculation happens to redescribe in an unnecessarily complicated way. That is a dispute about the word “new,” not about whether graph networks can rank formation energies — and it is exactly the distinction a headline number erases.

DeepMind has not published a structured, point-by-point reply to either critique as of this writing; the disagreement remains visible in the literature rather than resolved by it. What has happened instead is that the underlying programs kept running. The Materials Project, whose convex-hull data GNoME was trained on, continues to operate as an open institutional resource [7]. The A-Lab team has since built a second-generation platform capable of handling air-sensitive materials and applied it to a harder synthesis target — solid-state lithium conductors — with a published, quantified improvement in hit rate over the campaign, discussed below [13]. And in October 2025, Nature’s own news desk ran a retrospective assessment opening with almost the same framing this article opens with: “When the pioneering artificial intelligence (AI) firm Google DeepMind announced almost two years ago that it had used a deep-learning AI technique to discover 2.2 million new crystalline materials, it seemed to herald a thrilling new era of accelerated materials research,” before surveying the accumulated criticism and the continued, more cautious progress underneath it [4]. Two years on, in other words, the field’s own verdict is neither vindication nor retraction: it is an ongoing argument about what the pipeline has actually produced, conducted in public, by people with no stake in AI hype either way.

That argument belongs at the top of this article rather than at the bottom because it changes what every number after this point is allowed to mean. This is a piece about the industrialization of materials discovery — about turning a process humans used to run by trial, error, and inherited craft into something closer to a breeding program, with generation, selection, and culling happening at machine speed. But a breeding program is only as good as its selection step, and the specific, documented failure mode in both flagship 2023 demonstrations is that the selection step — the part where a candidate is checked against everything already known and everything a furnace can actually produce — was weaker than the generation step made it look. Keeping that failure visible, rather than letting “AI discovered millions of materials” harden into settled fact, is the discipline this article tries to hold from here on.

Materials have always evolved by the slowest variation-and-selection loop humanity runs

Materials science looks, from a long enough distance, like biological evolution running on an almost geological clock. A craftsperson varies an alloy’s composition, a firing temperature, a quenching method; most variants are indistinguishable or worse than what came before; a rare variant is noticeably better, gets copied, and becomes the new baseline against which the next round of variation is judged. Bronze displaced stone tools over centuries. Steel’s displacement of wrought and cast iron as the default structural metal, once the Bessemer process made it cheap, still took the better part of a generation to work through railroads, bridges, and buildings. Each transition is a lineage: a later alloy inherits most of its composition and processing history from the one before it, with a handful of new variations layered on top, exactly the way a new species inherits nearly all of its genome from its immediate ancestor.

The Materials Genome Initiative — the 2011 U.S. federal program that gave this article’s underlying subject its name — was founded on a blunt institutional admission of how slow that loop had become. Its founding strategic plan states the baseline directly: “the time frame for incorporating new classes of materials into applications is remarkably long, typically about 10 to 20 years from initial research to first use” [5]. Its worked example is the lithium-ion battery, chosen precisely because it is one of the faster cases on record: proposed as a laboratory concept in the mid-1970s, it reached wide market adoption only in the late 1990s — roughly 20 years — and the same 2011 document notes that “even now, 40 years later, lithium ion batteries have yet to be fully incorporated into the electric car industry” [5]. That is a textbook technology-adoption S-curve with an unusually long, still-unfinished tail: a slow initial climb through discovery and development, a long plateau while manufacturing, certification, and supply chains catch up, and — in this case — a second inflection decades later as electric vehicles finally pulled the same chemistry into a new market at new volume. The Initiative’s own seven-stage model of that curve — discovery, development, property optimization, systems design and integration, certification, manufacturing, and deployment — makes explicit what a headline “20 years” compresses: five of those seven stages have nothing to do with finding the material at all [5].

ADVERTISEMENT
A single furnace door swinging open on one crucible inside a long bank of otherwise closed furnaces in an automated synthesis lab
Figure 2. Even inside the fastest lab built for this, one material still bakes alone, on its own clock — the rate-limiting step materials science has run since the first kiln.Image prompt and art direction by Brecht Corbeel; generation pending.

Part of why that tail is so long is that cost and performance in most technologies fall with cumulative production rather than with elapsed time alone — the empirical pattern known as Wright’s Law, which a 2013 statistical survey of 62 technologies found to be the single best predictor of future cost decline among the competing models tested, tied to “learning by doing” rather than to the calendar [15]. Software and purely computational predictions can rack up cumulative “production” — training runs, candidate structures, simulated trials — at almost zero marginal cost. A physical material cannot: its learning curve only accumulates one furnace cycle, one pilot-plant batch, one qualification run at a time, which is precisely why compressing the computational front end of the pipeline does not automatically compress the tail behind it.

The National Research Council’s own estimate of the achievable prize, quoted approvingly in the Initiative’s founding document, was that integrating computational tools across that pipeline could shorten “the materials development cycle from its current 10-20 years to 2 or 3 years” [5] — and the Initiative’s own stated goal, more conservatively, was “a time reduction of greater than 50 percent” [5]. Fifteen years later, the Initiative’s own public page has not moved that number: it still states, in 2026, that “it can take 20 or more years to move a material after initial discovery to the market” [6]. That is worth sitting with before the rest of this article gets to the exciting part. The institution built specifically to compress this lag has spent a decade and a half stating the same 20-year figure it opened with, even as the tools available to attack the discovery end of the pipeline have changed completely. Either the promised compression is arriving on a longer clock than the discovery tools improved on, or it has arrived selectively — faster for some material classes than others — in a way a single round headline number cannot show. Both readings point the same direction: whatever has industrialized over the past decade, it has not yet been the whole seven-stage pipeline. It has been the first stage, and arguably just the first half of the first stage.

The industrialized pipeline generates candidates far faster than it used to, one piece at a time

What has genuinely changed since 2011 is the cost of generating a candidate in the first place, and it has changed in three distinct technological layers that are worth separating rather than treating as one undifferentiated “AI for materials” story.

The base layer is a shared ancestral database. The Materials Project, a Department of Energy effort based at Lawrence Berkeley National Laboratory, describes itself as “a decade-long effort… to pre-compute properties of ‘materials’ and make this data publicly available, with the intent of accelerating the process of materials discovery” [7]. Its method is density-functional theory applied at industrial scale: instead of a chemist calculating one compound’s properties by hand, a computing cluster calculates hundreds of thousands of them, all comparable on the same footing, all searchable. This is the “genome” in Materials Genome Initiative in a literal sense — a shared, standardized library of known and hypothesized structures against which any new candidate can be checked for both plausibility and lineage. Every dispute in the previous section runs through this layer: GNoME’s convex hull is built on Materials Project data, the A-Lab’s target list was drawn from Materials Project and DeepMind calculations together, and Cheetham and Seshadri’s own check-against-ICSD method is only possible because a comparably organized reference of known structures already existed to check against [1, 9, 2].

The second layer breeds new candidates rather than merely ranking existing ones. GNoME itself is a candidate-ranking and stability-prediction system: it does not invent structures from nothing so much as generate variations on known crystal families and filter them by predicted energy. A more recent tool, Microsoft Research’s MatterGen, is explicitly generative in the stronger sense — a diffusion model that builds a crystal structure atom type, coordinate, and lattice at a time, the way an image-diffusion model builds a picture from noise. Compared with earlier generative approaches, MatterGen’s own reported benchmark is that its structures are “more than twice as likely to be new and stable, and more than ten times closer to the local energy minimum” than prior methods, and the team went further than a purely computational demonstration: they fine-tuned the model to target a specific magnetic property, generated a candidate structure, synthesized it, and measured the resulting property “within 20% of our target” [8]. That is a genuinely different kind of evidence than a convex-hull count, because it closes the loop from proposal to physical confirmation on at least one example, inside the same paper, rather than deferring the check to a separate lab and a separate publication months or years later.

A powder-dosing station's hopper releasing a measured stream of precursor powder into an empty crucible on a scale, mid-pour
Figure 3. The recipe being followed here was written by a model, not a chemist — the step generative materials design has industrialized.Image prompt and art direction by Brecht Corbeel; generation pending.

The third layer is where the candidates meet a furnace, and it is the layer both 2023 critiques targeted. The A-Lab platform pairs a handful of robotic arms with a bank of furnaces and an automated diffraction bay: Berkeley Lab’s own account of the facility describes three robotic arms, eight furnaces, roughly 200 stocked precursor powders inside a 600-square-foot room, running 100 to 200 experiments a day — 50 to 100 times the throughput of a human researcher working the same bench by hand [10]. The lab’s principal investigator, Gerbrand Ceder, framed the urgency in blunt terms: “we need materials solutions for things like the climate crisis that we can build and deploy now” [10], and staff scientist Yan Zeng described the operating loop as letting the system “try something, analyze the data, and then decide what to do next” without waiting for a human to close each cycle [10]. Mechanically, the platform reads recipes proposed by a language model trained on the historical synthesis literature, dispenses precursor powders by robot, fires them, and reads the resulting X-ray diffraction pattern to judge success — 353 total experiments across the original 17-day campaign, according to the paper’s reported figures, refining recipes for compositions that failed on the first attempt through an active-learning loop grounded in reaction thermodynamics [9]. That is precisely the loop Leeman and colleagues checked and found wanting at the interpretation step, not the mechanization step [11] — the robots ran the reactions and read the patterns exactly as designed; the dispute is about what the patterns were then read to mean.

ADVERTISEMENT
Two furnace doors just opened together on a freshly fired batch, a robot arm reaching toward the nearer crucible before the tray has fully cooled
Figure 4. This is the moment a claimed discovery is either confirmed or quietly deflated — and, campaign after campaign, where most claimed candidates fail.Image prompt and art direction by Brecht Corbeel; generation pending.

None of this is confined to one lab. The Acceleration Consortium at the University of Toronto now coordinates more than 30 self-driving labs worldwide, describing their combined effect as a 10-to-100-fold improvement in both the speed and the cost of materials and molecular discovery, across a membership of more than 200 organizations and 60 industry partners spanning health, energy, construction, and electronics [12]. This is the point at which the evolutionary metaphor stops being decorative: a self-driving lab is not a single specialized instrument but closer to a genus — one underlying design pattern (propose, synthesize, measure, retrain) radiating into dozens of independently operated implementations, each adapted to its own chemistry the way a founder population radiates into distinct species once it disperses into separated niches. The Berkeley A-Lab and its 2026 successor for air-sensitive lithium compounds, discussed next, are two members of that genus rather than two unrelated inventions [13].

Once search is cheap, selection moves downstream — to synthesis, scale, and supply chains

The evolutionary framing does real analytic work once the pipeline above is treated as a selection funnel rather than a single machine. A useful way to write down what a candidate has to survive to become a deployed material is a chain of conditional probabilities, each stage filtering out most of what entered it:

P(deployed)=P(stable)×P(synthesizablestable)×P(scalablesynthesizable)×P(supply-chain viablescalable) P(\text{deployed}) = P(\text{stable}) \times P(\text{synthesizable} \mid \text{stable}) \times P(\text{scalable} \mid \text{synthesizable}) \times P(\text{supply-chain viable} \mid \text{scalable})

GNoME and MatterGen attack only the first term, and the 2023-2024 dispute is essentially an argument about whether even that first term was measured honestly. But the more interesting evolutionary claim in this article’s premise is what happens to the later terms once the first one becomes cheap to compute at scale: selection pressure does not disappear, it relocates to whichever stage is now the tightest bottleneck — the way a population that solves one predator pressure simply comes under stronger selection from the next one, be it food scarcity or a different predator.

Two documented cases show that relocation happening in real time. The first is a straightforwardly quantitative funnel. In a 2022 Science paper, Ziyuan Rao and sixteen co-authors used an active-learning loop combining machine learning, density-functional theory, thermodynamic calculation, and physical experiment to search the compositional space of high-entropy “Invar” alloys — a space with millions of possible compositions — for combinations with an unusually low thermal expansion coefficient [14]. Of that space, the team physically synthesized and characterized only 17 alloys; of those 17, two showed the target property, with thermal expansion coefficients around 2×10⁻⁶ per kelvin at 300 kelvin [14]. Millions of candidates in, two winners out, with the entire discipline of the project spent deciding which 17 of the millions were worth the furnace time — synthesis capacity, not compositional imagination, was the visible constraint the whole time.

A long rack of loaded crucible trays queued behind a single furnace slot, one tray's carrier caught mid-lift toward the only open door
Figure 5. Millions of candidates, dozens synthesized, a handful that survive the next test: the funnel that decides which predictions become materials.Image prompt and art direction by Brecht Corbeel; generation pending.

The second case shows the selection funnel getting measurably more efficient across a single campaign, which is closer to artificial selection intensifying generation over generation than to one static filter. A 2026 paper from the Ceder group describes an upgraded A-Lab, built specifically to handle air-sensitive materials, applied to searching for lithium halide spinel conductors — a candidate family of solid-state battery electrolytes, precisely the application GNoME’s own abstract had flagged as a target use case for its predictions [1, 13]. Across 352 total synthesized samples exploring pairwise combinations among 19 metals, an agentic reasoning layer directing the campaign — alternating what the authors describe as “abductive” reasoning to interrogate anomalies in regions already explored and “inductive” reasoning to expand into unvisited chemical space — lifted the fraction of compositions meeting both a minimum ionic-conductivity threshold and high phase purity from 1.33% in the first 75 agent-proposed samples to 5.33% in the final 75 [13]. That fourfold improvement in hit rate is the selection process itself evolving mid-campaign, in the same sense that a breeder’s eye for a trait sharpens across successive generations of a breeding program even while the underlying genetic variation it draws on stays the same.

Both cases also show the fitness function itself moving downstream, which is the more consequential evolutionary point. MatterGen’s authors do not stop at thermodynamic stability: they explicitly demonstrate fine-tuning the model to jointly target a physical property — magnetic density — and a chemical composition with low supply-chain risk, meaning the training signal itself now encodes an economic and geopolitical constraint alongside a thermodynamic one [8]. That is the modern equivalent of moving from viability selection — can the organism survive at all — to fecundity and dispersal selection — can it actually propagate and spread once it has survived. A material that is stable, synthesizable, and even scalable can still fail this later stage if its synthesis depends on an element concentrated in one politically unstable jurisdiction, or a processing step no existing factory is tooled for; incumbency itself becomes a selective force, in the sense that an existing material’s installed manufacturing base, certification history, and supply relationships function as a lock-in that a marginally superior newcomer has to out-compete on more than raw performance to displace. Steel did not win against iron on strength alone; it won because the Bessemer process made steel’s manufacturing economics superior at exactly the moment structural demand was exploding. The same logic now runs inside a training objective rather than a boardroom, but it has not gone away.

Whether the discovery-to-deployment lag actually compresses before 2035 is a falsifiable question, not a foregone conclusion

Everything documented above is a claim about the discovery end of the seven-stage pipeline getting faster. Whether that translates into the Materials Genome Initiative’s original goal — a 20-year lag compressed toward the National Research Council’s 2-to-3-year aspiration [5] — is a separate, forward-looking claim, and it deserves the same discipline the rest of this series applies to any statement about the future: a stated horizon, stated assumptions, a named indicator, and an explicit way to be proven wrong.

The claim, and whose claim it is. This article’s own prediction, on a horizon to 2035: for at least one narrow material class — plausibly a solid-state battery electrolyte, given how much of the pipeline above is already pointed at that target [1, 13] — the interval between a computational prediction and a commercially deployed product will fall to under a decade, breaking meaningfully below the Materials Genome Initiative’s own 10-to-20-year historical baseline [5] for that one class, even if the broader distribution of material classes does not move nearly as fast.

Assumptions. This prediction assumes three things that are each independently uncertain. First, that the synthesis-stage disputes documented in this article’s opening section get resolved in the technical, boring way rather than the embarrassing way — that automated Rietveld refinement and disorder-aware structure prediction improve enough that a claimed “new material” reliably survives adversarial scrutiny of the kind Cheetham, Seshadri, Leeman, and their co-authors already applied once [2, 11]. Second, that at least one self-driving lab, among the more than 30 the Acceleration Consortium already coordinates [12], carries a specific candidate all the way through scale-up and qualification rather than stopping at a bench-scale demonstration, which is a commitment of capital and time this article has found no fully documented case of yet. Third, that the “10-100x faster and cheaper” figure the Acceleration Consortium reports for discovery [12] is not entirely absorbed by a correspondingly longer downstream qualification process — the risk this article’s second section already flagged, where the Materials Genome Initiative’s own 20-year figure has not moved in fifteen years of tool improvement [6].

Indicators. A reader checking this prediction in 2035 should look for a named material — not a named model, not a named database — with a publicly documented, unbroken chain from a specific computational or generative prediction, through a specific autonomous or semi-autonomous synthesis campaign, through pilot-scale manufacturing, to a specific commercial product incorporating it, with each step attributable to a dated publication or filing rather than a marketing claim. As of this writing in August 2026, this article has not found a case that clears that bar: the closest candidates — GNoME’s flagged solid-electrolyte and lithium-ion-conductor candidates [1, 3], MatterGen’s single synthesized proof-of-concept [8], and the upgraded A-Lab’s lithium halide spinel campaign [13] — are all still at or before the synthesis-and-characterization stage, several stages short of deployment on the Initiative’s own seven-stage model [5].

A crucible carousel fully loaded and still turning past a furnace bank where only one door is cycling, most furnaces standing idle
Figure 6. The falsifiable question for the next decade: does furnace capacity ever catch the carousel, or does the backlog just get longer.Image prompt and art direction by Brecht Corbeel; generation pending.

Disconfirmation. This prediction is falsified if, by the end of 2035, no material class shows a documented discovery-to-deployment interval under ten years, and it is more strongly disconfirmed if the aggregate pattern instead resembles lithium-ion’s own history — a genuinely fast laboratory-to-first-use step followed by a decades-long, still-unfinished tail before the material reaches its largest addressable market [5]. The evidence available today is compatible with either outcome. The generation side of the funnel has unambiguously industrialized: millions of candidates a year is now routine, whatever the argument about how many of them are genuinely new [1, 2]. The synthesis side has industrialized much further than a human bench chemist but is still measured in hundreds of samples per campaign, not millions [9, 14, 13]. And the stages beyond synthesis — certification, manufacturing tooling, supply-chain qualification — are not obviously touched by any of the technology described in this article at all. A prediction that ignores that gap is not extending the evidence; it is skipping past the part of the pipeline nobody has yet shown how to automate.

The artisanal loop has closed into an industrial breeding program, but synthesis is still the bottleneck the critics identified

Put the pieces back together and the framing in this article’s title earns itself rather than just describing a metaphor. Materials development was always a variation-and-selection process; it simply used to run at the pace of individual careers, with each generation of alloys or ceramics inheriting most of its properties from the last and adding a handful of new variations discovered by trial, accident, or craft intuition passed down as tacit knowledge. The Materials Genome Initiative’s founding insight in 2011 was that this loop could be industrialized the way agriculture industrialized selective breeding — not by inventing a new mechanism, but by running the same mechanism at radically higher throughput, with a shared genome of computed structures replacing tacit craft knowledge as the thing each new generation of candidates is built from [5, 7].

Fifteen years later, one side of that industrialization is real and well documented: generating candidates and ranking them by predicted stability has gone from a hand calculation to a background process producing millions of outputs a year [1, 8]. But the 2023 disputes this article opened with are not a footnote to that success; they are the discovery, made in public and under adversarial scrutiny, that the selection step had not kept pace with the generation step — that a large fraction of “new” predicted structures were known compounds in a disorder-blind disguise, and that a flagship autonomous-synthesis demonstration built on those predictions may, on the critics’ own count, have produced no genuinely new material at all [2, 11]. The honest reading of the evidence assembled here is that this is exactly what an evolutionary process looks like when mutation gets radically cheaper while selection does not: an explosion of variation with no corresponding explosion in verified fitness, until the selection apparatus — better disorder modeling, better automated diffraction interpretation, more furnace and characterization capacity — catches up.

Whether that catch-up happens on the multi-year timeline the self-driving-lab community is betting capital on, or on the multi-decade timeline the Materials Genome Initiative’s own unmoved 20-year figure suggests is more typical, is the falsifiable question this article has tried to leave a reader equipped to actually check [12, 6]. The breeding program is real. So, on the evidence collected by the field’s own harshest internal critics, is the fact that almost everything claimed as bred so far still has to survive being born.