Where the roads actually diverge by 2035
Ask whether artificial intelligence will touch science and medicine by 2035 and the question answers itself: it already has. A protein-structure predictor won its creators a share of a Nobel Prize in 2024. A single deep-learning system has proposed millions of candidate inorganic crystal structures. A GPT-4-driven agent has planned and executed real palladium-catalysed chemistry with no human in the physical loop. None of that is speculation; all of it is dated, published, and — this is the part usually skipped — already contested in places.
The more interesting question, and the one this article is built to answer, is not whether but which kind. There is more than one coherent way the current trajectory can resolve by 2035, and the differences between them are not cosmetic. This article names four: Autonomous Discovery, in which AI systems increasingly originate the hypotheses that drive real findings rather than merely testing ones humans propose; Validation Bottleneck Persists, in which computational throughput keeps compounding while wet-lab synthesis and clinical trials remain the rate-limiting step regardless of how fast the models get; Regulatory Adaptation, in which frameworks built for static software genuinely retool themselves for continuously updating AI; and Reproducibility Crisis Deepens, in which the sheer volume of AI-assisted scientific claims outruns the field’s capacity to check them, and trust erodes as a result.
Throughout, five kinds of statement are kept visibly separate. A fact is something disclosed in a peer-reviewed paper, a regulatory record, or an institutional report. A vendor claim is a company’s own statement about its own system, reported as a claim because the company has a stake in the answer. Analysis works out a consequence of stated facts. A scenario is one internally consistent way the future could go, presented beside its alternatives rather than as the likely one. Each scenario below carries an explicit set of assumptions, names the observable indicators that would show the world moving toward it, and states in advance what would disconfirm it.
The documented present
The clearest achievement in this field is also the one furthest from any live controversy. In 2021, DeepMind’s AlphaFold system became the first computational method to regularly predict protein structures with atomic accuracy even for sequences with no similar structure already known, reaching a median backbone accuracy of 0.96 Å (root-mean-square deviation) against the next-best method’s 2.8 Å at the 2020 Critical Assessment of protein Structure Prediction, and an all-atom accuracy of 1.5 Å against 3.5 Å [1]. The paper is also explicit about where the method still struggles: accuracy drops substantially once the available multiple sequence alignment has fewer than roughly 30 related sequences to draw on, and the model performs worse on proteins whose contacts are mostly between separate chains rather than within one — precisely the bridging domains that hold large protein complexes together [1]. In 2024, the Royal Swedish Academy of Sciences awarded half of the Nobel Prize in Chemistry jointly to Demis Hassabis and John Jumper for protein structure prediction, and the other half to David Baker for computational protein design; by the time of the award, the resulting AlphaFold Database had grown from roughly 360,000 predicted structures at its 2021 launch to some 200 million, used by more than a million researchers in nearly every country with a research sector [2].
Weather forecasting supplies a third, less contested example of the same underlying pattern — genuine scientific throughput rather than a demo. Lam and colleagues’ GraphCast, trained directly on historical reanalysis data rather than on the physics-based numerical models meteorology has relied on for decades, predicts hundreds of atmospheric variables ten days ahead at 0.25-degree global resolution in under a minute, and the peer-reviewed evaluation found it more accurate than the world’s leading operational deterministic forecasting system on 90% of 1,380 verification targets, including measurable improvements in tracking severe events such as tropical cyclones and atmospheric rivers [9]. Unlike GNoME or the A-Lab below, GraphCast’s claims rest on a forecasting task with an unambiguous, fast, and repeated real-world scoring mechanism — the weather itself arrives within days and settles every dispute about whether a given forecast was right — which is precisely the structural feature the more contested cases in this section lack.
Materials science offers a structurally similar story with a much rougher edge. In November 2023, Google DeepMind published GNoME, a deep-learning system that screened candidate crystal structures at a scale no prior method approached: the peer-reviewed paper describes searching 2.2 million candidate structures and identifying roughly 381,000 that its models computed to be thermodynamically stable [3]. DeepMind’s own announcement of the same work states, as a vendor claim rather than an independently audited figure, that this expanded the total library of experimentally or computationally known stable inorganic materials from around 48,000 to roughly 421,000, that the system’s precision at predicting stability rose to about 80% from roughly 50% for prior methods, and that 736 of GNoME’s predicted structures were subsequently found to match materials independently synthesized and reported by outside laboratories [4]. That is a genuinely large, dated, checkable claim. It is also one that two materials chemists at UC Santa Barbara, Anthony Cheetham and Ram Seshadri, examined directly in a peer-reviewed 2024 perspective and found wanting: writing that they found “scant evidence for compounds that fulfill the trifecta of novelty, credibility, and utility,” and noting that GNoME’s outputs are exclusively crystalline inorganic compounds rather than the broader universe of polymers, glasses, metal-organic frameworks, and composites that “materials” implies to most working scientists [5].
A parallel and even sharper case sits one step closer to the physical world. Also published in Nature in November 2023, the A-Lab at Lawrence Berkeley National Laboratory combined ab initio phase-stability data, literature-trained synthesis-recipe extraction, and a robotic platform running three linked stations — sample preparation, heating, and X-ray diffraction characterization — to run 353 experiments across 17 days of largely unattended operation [6]. Independent chemists subsequently challenged the paper’s diffraction analysis directly, arguing in a critique that the AI-driven Rietveld refinement used to identify which compounds had actually formed was unreliable and that many of the claimed products were, on manual re-examination, ordered versions of compounds already known to be disordered [8]. Nature issued a formal author correction in January 2026: after manually re-analysing the diffraction data, the original authors reported that the prediction platform had reached the correct conclusion in 36 of 40 previously reported successes, with four inconclusive, clarified that the paper’s claim of “novel” materials meant new to the prediction platform’s own training data rather than new to science, and removed one compound that had mistakenly been included in that training set to begin with [7]. The current peer-reviewed record for the paper describes 36 of 57 attempted targets successfully synthesized [6]. A training-data contamination bug, caught by outside critics roughly two years after publication rather than by the original team — inside one of the most widely reported “autonomous discovery” demonstrations of the decade — is not a hypothetical failure mode. It already happened.
Two further, more encouraging data points belong in this same documented present, because they show the loop closing rather than breaking. Boiko and colleagues’ Coscientist, a GPT-4-driven system augmented with internet search, documentation search, code execution, and laboratory automation, was demonstrated performing six distinct chemistry tasks including the successful reaction optimization of a palladium-catalysed cross-coupling, with the authors devoting explicit discussion to the system’s safety implications and to measures for preventing its misuse [10]. And in 2025, Google’s “AI co-scientist,” a multi-agent system built on Gemini and organised around self-play scientific debate and tournament-based hypothesis ranking, produced drug-repurposing candidates for acute myeloid leukemia that were confirmed to inhibit tumour viability at clinically relevant concentrations across multiple cell lines in wet-lab follow-up, and identified epigenetic targets for liver fibrosis that showed statistically significant anti-fibrotic activity in human hepatic organoid experiments [11]. The same system also independently proposed a mechanism for how capsid-forming phage-inducible chromosomal islands spread across bacterial species — a hypothesis that matched unpublished experimental results already sitting in a collaborating laboratory, arrived at within days rather than the years the original wet-lab programme had taken [11]. Its authors are explicit that the system is intended as “a collaborative tool for scientists” designed to “augment human ingenuity,” not replace the humans who frame the research question, and that every validation experiment in the paper proceeded under “expert-in-the-loop guidance” [11].
That is the documented present in miniature: real breakthroughs, real scale, real autonomous execution — sitting directly beside a genuine, dated, embarrassing correction inside one of the field’s own flagship demonstrations. The four scenarios below are four different ways that tension can resolve.
Scenario A: Autonomous Discovery
Definition. By 2035, AI systems function as genuine co-discoverers across a widening set of scientific domains: they originate hypotheses, design the experiments that test them, and increasingly close the loop with minimal human hypothesis-generation, while humans shift toward framing research questions at a higher level and validating conclusions rather than generating candidate ideas themselves.
Evidence baseline. The strongest currently available evidence for this scenario is the AI co-scientist’s wet-lab-confirmed drug-repurposing and liver-fibrosis findings and Coscientist’s autonomous, end-to-end execution of real chemistry [11, 10]. Both are dated, peer-reviewed or peer-adjacent, and independently checkable.
Assumptions. This scenario requires that today’s narrow successes — each still explicitly framed by a human research question and validated under expert supervision — generalise outward without a proportional rise in the human framing effort required per discovery. It also requires that the rate of genuinely novel, rather than literature-recapitulating, validated hypotheses produced by such systems keeps rising, rather than plateauing once the easiest recombination of existing knowledge has been exhausted.
Observable indicators. Look for published discovery-attribution studies that credit an AI system as the primary originator of a hypothesis rather than a tool assisting a human-originated one; a falling ratio of expert-hours to validated discovery as agentic systems mature; and a rising share of AI-originated findings that are, on independent scrutiny, actually new to the field rather than rediscoveries of results already sitting unpublished in someone’s laboratory, as the phage-inducible-chromosomal-island case turned out to be [11].
Disconfirmation condition. This scenario is disconfirmed if independent audits repeatedly find that the fraction of “AI-discovered” results that are genuinely novel to the published literature — rather than restatements, minor extensions, or rediscoveries of unpublished prior work — stays flat or falls, echoing what happened on closer inspection to GNoME and the A-Lab’s headline claims [5, 8]; or if expert-hours required per validated discovery do not measurably decline as these systems mature.
Scenario B: Validation Bottleneck Persists
Definition. Computational and generative throughput keeps compounding, but wet-lab synthesis, materials characterization, and clinical trials remain the binding constraint on how much of that output ever becomes a checked fact — and the gap between candidates generated and candidates confirmed widens rather than closes by 2035.
Evidence baseline. The clearest number here is a ratio, not a rate: GNoME’s own account describes 2.2 million candidate structures screened against 736 that outside laboratories are reported to have independently synthesized and matched [3, 4] — a confirmation rate on the order of one in three thousand, even taking the vendor’s own figures at face value. The physical ceiling on the other side of that ratio is visible in the A-Lab’s own operating record: 353 experiments and 36 to 41 confirmed compounds across 17 days of fully robotic, round-the-clock operation [6, 7] — a small number of confirmed materials per day even at the frontier of automation. On the medical side, Wu and colleagues’ analysis of FDA-cleared AI devices found that the evaluation process itself can mask vulnerabilities that only appear once a device is deployed on real patients, and argued that no established best practices existed, comparable to those for academic clinical trials, for evaluating commercially available algorithms before their approval [15] — a description of a system built around a single retrospective evaluation snapshot rather than continuous confirmation.
Assumptions. This scenario assumes that robotic-lab throughput and clinical-trial capacity do not scale the way compute has, because synthesis, characterization, and regulated human-subjects research all sit on top of physical, labour, and ethical constraints that a faster model does not relax. It assumes the confirmation side of the ledger is a comparatively hard floor.
Observable indicators. Track the ratio of computationally flagged candidates to independently confirmed ones over time in materials science and drug discovery; track autonomous-lab throughput in confirmed compounds per day as a direct measure of whether automation is closing or merely relocating the bottleneck; and track median clinical-trial cycle times for AI-flagged drug candidates against the historical baseline.
Disconfirmation condition. This scenario is disconfirmed if autonomous-lab throughput scales by an order of magnitude or more — thousands rather than tens of confirmed compounds per year from a comparable facility — or if clinical-trial cycle times fall sharply because of AI-native adaptive trial design, closing rather than widening the gap between what is computed and what is confirmed.
Scenario C: Regulatory Adaptation
Definition. By 2035, medical-device regulatory frameworks have successfully retooled themselves for AI systems that update continuously rather than shipping as fixed software, with mechanisms like the FDA’s predetermined change control plan becoming the default expectation for a new AI-enabled device rather than a novel exception, and extending well beyond today’s radiology-heavy device mix into higher-stakes, more dynamic clinical domains.
Evidence baseline. On 3 December 2024, the FDA finalized guidance recommending that a predetermined change control plan (PCCP) describe, in advance, the modifications a manufacturer intends to make to an AI-enabled device, the methodology for developing and validating those modifications — including bias-mitigation strategies and post-market surveillance — and an assessment of their impact, so that pre-authorized updates would not each require a fresh marketing submission; the final guidance broadened the scope from the 2023 draft’s machine-learning-only focus to cover all AI-enabled device software functions [13]. That sits alongside a genuinely fast-growing base to adapt: the FDA’s own list shows 1,451 AI-enabled medical devices authorized since it began tracking such devices in 1995 through the end of 2025, with 295 authorized in 2025 alone [14].
Analysis. The same dataset that shows adaptation happening also shows how narrow its base still is. Radiology accounted for 1,104 of those 1,451 devices — 76% of the total — with the share by year running 80% in 2023, 73% in 2024, and 75% in 2025 [14]. Radiology AI is a comparatively well-bounded pattern-recognition task evaluated against a fixed set of image-derived outcomes; it is not obviously representative of continuously learning systems making dosing, triage, or diagnostic decisions in faster-moving, higher-stakes settings. And Wu and colleagues’ finding that most existing approvals rest on a single retrospective evaluation snapshot rather than ongoing real-world monitoring [15] describes an apparatus still largely built, in practice, for static software — even as the PCCP pathway now exists on paper for the dynamic case.
Assumptions. This scenario assumes the PCCP framework, so far mostly untested against systems with a genuinely dynamic risk profile, generalises to higher-stakes domains without a safety incident severe enough to trigger a policy reversal, and that other major regulators converge on broadly comparable mechanisms rather than fragmenting into incompatible regimes.
Observable indicators. Watch the count of devices actually operating under an FDA-approved PCCP with a disclosed modification history, rather than merely holding one on file; the growth of AI device authorizations outside radiology as a share of the yearly total; and whether comparable continuous-update pathways emerge in other major regulatory jurisdictions.
Disconfirmation condition. This scenario is disconfirmed if, by the horizon, PCCP adoption stays marginal — most AI-enabled devices still cleared and updated through the older, single-submission pathway — or if a high-profile AI-device safety failure triggers a retreat toward stricter, case-by-case resubmission requirements rather than continued expansion of continuous-update mechanisms.
Scenario D: Reproducibility Crisis Deepens
Definition. The volume of AI-assisted scientific claims keeps outrunning the field’s capacity to audit them. Leakage-driven and validation-driven corrections keep surfacing well after publication and press coverage, at an accelerating rather than a slowing rate, and trust in AI-flagged results generally — not just the specific ones later corrected — erodes as a consequence.
Evidence baseline. The general baseline predates and is independent of any single AI-for-science claim: Kapoor and Narayanan’s systematic survey found data leakage affecting 329 papers across at least 17 research fields that had adopted machine-learning methods, in some cases producing conclusions the authors describe as “wildly overoptimistic,” and built an eight-part taxonomy of leakage types ranging from textbook errors to genuinely unresolved research problems; as one worked example, they show that a set of published papers claiming complex ML models outperformed decades-old logistic regression for predicting civil conflict failed to reproduce once the leakage in their evaluation setup was corrected, with the simpler statistical models actually performing comparably all along [12]. That general pattern has a startlingly concrete specific instance: the A-Lab’s own author correction discloses that one compound had been mistakenly included in the platform’s own training data — precisely the kind of leakage Kapoor and Narayanan’s taxonomy describes — inside a paper that had already been celebrated in hundreds of press outlets as a landmark demonstration of autonomous scientific discovery, and the error was caught by outside chemists roughly two years after publication, not by the original team [7, 8]. GNoME drew an analogous, if less severe, correction from domain experts who, despite the scale of the initial announcement, found “scant evidence” that the predicted materials met ordinary standards of novelty, credibility, and utility [5].
Assumptions. This scenario assumes publication and media incentives continue to reward scale and speed claims faster than the community’s replication and auditing capacity grows, and that standardized leakage-preventing reporting practices — such as the “model info sheets” Kapoor and Narayanan propose [12] — remain optional recommendations rather than enforced norms at major venues.
Observable indicators. Track the rate of published corrections and critique papers targeting flagship AI-for-science claims specifically, as distinct from the general background rate of scientific corrections; track whether journals and funders actually require standardized leakage-preventing reporting rather than merely recommending it; and track independent replication rates for AI-flagged findings specifically, set against the replication baseline for traditionally derived findings in the same fields.
Disconfirmation condition. This scenario is disconfirmed if the correction and critique rate targeting flagship AI-for-science claims declines as standardized reporting spreads, and if independent re-analyses of celebrated results increasingly confirm rather than walk back the original headline claims, the way the A-Lab and GNoME cases did not.
What the four scenarios share, and where they actually pull against each other
None of these four is presented as the likely one, and they are not mutually exclusive — the documented present already contains real elements of all four simultaneously. That is worth dwelling on rather than resolving away. The AI co-scientist’s wet-lab-confirmed findings are genuine evidence for Autonomous Discovery, and the same paper’s own description of “expert-in-the-loop guidance” at every validation step is simultaneously evidence that Validation Bottleneck Persists is already partly true [11]. The FDA’s PCCP guidance is real evidence for Regulatory Adaptation, and Wu and colleagues’ finding that most approvals still rest on a single retrospective snapshot is evidence that the older, static regime it is meant to replace remains the norm in practice [13, 15]. The real 2035 question is less “which one happens” than which of these four tendencies becomes the frame the field organises itself around — which is to say, which one wins enough attention, funding, and institutional weight to set the terms for the other three.
There is also a structural tension worth naming explicitly between two of the four, because it is not obvious and it matters for reading the other three. Regulatory Adaptation, as a scenario, depends on regulators being able to trust that a continuously updating AI system’s behaviour is auditable enough to pre-authorize its own future changes with reasonable confidence. Reproducibility Crisis Deepens, as a scenario, is precisely the claim that the field’s auditing capacity is not keeping pace with the claims it needs to check. A regulator confident enough to hand a manufacturer a predetermined change control plan is making a bet that looks a great deal like the bet Nature’s editors and reviewers made about the A-Lab paper before an outside critique and a two-year-later correction revealed a training-data contamination bug the original review process had missed [7]. If Reproducibility Crisis Deepens turns out to be the dominant pattern, Regulatory Adaptation becomes harder to sustain for exactly that reason — not because the regulation itself is flawed, but because the evidentiary base it depends on is exactly the kind of AI-generated claim this article’s evidence shows is currently under-audited.
What to take away
Four things are documented fact at the point this article was written, and none of them is a prediction. AlphaFold’s structure-prediction accuracy is real, independently verified across a fourteenth-round international assessment, and recognised with a Nobel Prize [1, 2]. GNoME’s candidate-generation scale is real, and so is the peer-reviewed expert critique that found most of its output falling short of ordinary standards of novelty, credibility, and utility [3, 5]. Autonomous laboratories and research agents have executed and reported genuine science with reduced human involvement in the physical loop, and one of the most celebrated examples of that same category required a formal published correction after outside chemists caught an error the original review process missed [10, 11, 7]. And a regulator has already built, and is beginning to use, a real mechanism for pre-authorizing a continuously updating AI system’s future changes, even though the devices it has approved so far skew heavily toward one comparatively well-bounded specialty [13, 14].
Everything beyond those four facts in this article is either a company’s own claim about its own system, an analysis of what a stated fact implies, or one of four named scenarios carrying an explicit horizon, explicit assumptions, and an explicit way to be proven wrong. The discipline this article asks of a reader returning to it as 2035 approaches is the same one it has tried to hold to in writing it: check each scenario’s stated indicators against what actually happened, and notice which of its assumptions turned out to be the one that mattered.