Five ways a promising result quietly stops being true

Most AI-for-science projects do not fail because the underlying model is weak. They fail because of five ordinary engineering decisions made early and rarely revisited: how the data was divided before training started, whether the loop connecting predictions to real experiments actually closes, what evidence a regulator will accept, whether a result can be reproduced on demand, and whether anyone checked a vendor’s number before building on it. All five are common enough to have generated their own peer-reviewed failure literature, and all five are fixable with disciplined engineering rather than a better architecture.

This is a working guide to those five decisions, aimed at someone doing the work — designing the split, running the loop, writing the submission, versioning the pipeline, or evaluating the tool. It stays in that register throughout: concrete practice, not conceptual survey.

The split is the experiment

In most of applied machine learning, a random train/test split is a reasonable default. In computational biology and chemistry, it is frequently the single decision that invalidates everything downstream, because the units in these datasets are rarely independent and identically distributed the way a random split assumes.

ADVERTISEMENT

The clearest documented case is a 2018 exchange in Science. Ahneman and colleagues had reported a model predicting C–N cross-coupling reaction yields from atomic, electronic and vibrational descriptors, evaluated with a random split. Chuang and Keiser’s comment showed that the experimental design could not actually distinguish a model trained on real chemical descriptors from one trained on randomly generated features of the same shape, in both retrospective and prospective test scenarios — the reported accuracy said less about the chemistry than it appeared to [3]. The failure was not a bad model. It was an evaluation that could not detect whether the model had learned chemistry or had simply exploited structure in how the random split happened to fall.

That single case generalises further than one reaction-yield paper. Kapoor and Narayanan surveyed the peer-reviewed literature across seventeen fields that have adopted machine learning and found leakage-driven reproducibility failures in 329 papers, spanning a taxonomy of eight distinct leakage types — from textbook errors like preprocessing on the full dataset before splitting, to subtler failures like temporal leakage, duplicate or near-duplicate records across the train and test sets, and non-independence between the sampling unit and the evaluation unit [1]. Their central methodological point is that leakage is not a coding bug to be caught by careful engineers; it is a property of the relationship between how data was generated and how it was split, and it recurs precisely because each field’s data has its own hidden dependency structure that a generic random split does not respect.

For molecular data specifically, the dependency structure is scaffold similarity: two compounds that share a chemical scaffold are not independent draws, and a model can often achieve high held-out accuracy by memorising a scaffold’s typical activity rather than learning a transferable structure–property relationship. For protein or genomic data, the dependency is often temporal or genealogical: sequences deposited close together in time, or proteins from the same family, share information that a naive random split leaks across the train/test boundary. For any data tied to physical instruments or acquisition batches, batch effects create a third dependency structure entirely — a model can learn to recognise the instrument or the day, not the biology.

The practical fix in each case is to split along the axis of the dependency, not at random:

Scaffold or cluster splitting. Group molecules by scaffold (or by a clustering of molecular fingerprints) before splitting, and assign entire groups — never individual members of a group — to train, validation or test. A held-out set built this way tests generalisation to new chemical space rather than interpolation within a chemical family already represented in training.

ADVERTISEMENT

Temporal splitting. Where data accumulates over time, train only on data available before a cutoff date and test only on data collected after it. This is the only split that matches the deployment condition a model will actually face — predicting the future from the past — and it is the split most often skipped because it is inconvenient: it throws away the most recent, often highest-quality data from training.

Blind, externally adjudicated evaluation. The Critical Assessment of protein Structure Prediction (CASP), the biennial blind trial that has evaluated protein structure prediction methods since 1994, offers a template worth borrowing even outside structural biology. Structures under active experimental determination but not yet public are used as prediction targets; participants submit predictions before the true structure is released, so the “test set” is not just held out but does not yet exist in any form a model could have leaked from. AlphaFold’s 2020 CASP14 result — accuracy competitive with experimental methods on the majority of targets — carried the weight it did specifically because the evaluation protocol made leakage structurally impossible, not merely unlikely [4]. A practitioner without access to a standing blind-trial infrastructure can approximate the same guarantee with a prospective holdout: freeze the model, then evaluate only on data that does not yet exist at freeze time.

Negative controls on the split itself. Before trusting a held-out score, run the identical pipeline with the labels randomly permuted (a y-scrambling control) and, separately, with the input features replaced by random noise of the same shape, exactly as Chuang and Keiser did to the cross-coupling model. If either control achieves a suspiciously high score, the split — not the model — is the thing to fix.

None of this is exotic statistics. It is closer to bookkeeping discipline: before any modelling begins, decide what axis of dependency exists in this specific dataset, split along it, and then verify the split with a control designed to fail if leakage is present.

Designing an active-learning loop that actually closes

A model that only ever predicts is not doing science; it is producing a ranked list. The loop closes only when a prediction leads to a real measurement, and that measurement changes what the model predicts next. Two recent, fully documented systems show what a closed loop looks like in practice, at opposite ends of the automation spectrum.

Burger and colleagues built a mobile robot that searched a ten-variable formulation space for a photocatalyst, guided by a batched Bayesian optimisation algorithm. Operating autonomously over eight days, it ran 688 experiments and converged on photocatalyst mixtures six times more active than the best in the initial candidate set [6]. Szymanski and colleagues built A-Lab, a system that combined literature-derived and computed candidate materials, an active-learning algorithm called ARROWS3 that revises predicted synthesis routes using each experiment’s outcome, and a robotic solid-state synthesis line. Over seventeen days of largely unattended operation it attempted 57 novel target compounds and successfully synthesised 36 of them — a 63 percent success rate — with the authors reporting that refinements to the active-learning component alone could raise that rate further on the same target set [5].

ADVERTISEMENT

Both systems share a structure worth naming explicitly, because it is what separates a genuine active-learning loop from a model that periodically gets consulted:

The acquisition function must trade off predicted value against uncertainty, not just rank by predicted value. A pure exploitation policy — always synthesise the top-ranked candidate — collapses quickly onto whatever region of chemical or materials space the initial training data already covered well, and stops finding anything the model did not already believe. The standard fix is an acquisition function such as expected improvement, which for a candidate xx↗ with predictive mean μ(x)\mu(x)↗, predictive standard deviation σ(x)\sigma(x)↗, and current best observed value f+f^{+}↗ is

EI(x)=E ⁣[max⁡(f(x)−f+, 0)], \mathrm{EI}(x) = \mathbb{E}\!\left[\max(f(x) - f^{+},\ 0)\right], ↗

a quantity that rewards both a high predicted mean and high predictive uncertainty. In practice this means the loop deliberately spends some experimental budget on candidates the model is unsure about, not only on candidates it is confident are good — because the uncertain ones are where each experiment teaches the model the most.

Batch size interacts with the value of parallel experiments in a way worth costing out explicitly. If a single proposed experiment succeeds with probability pp↗, and a batch of kk↗ proposals were independent, the probability that at least one succeeds is

P(at least one success in k)=1−(1−p)k, P(\text{at least one success in } k) = 1 - (1-p)^{k}, ↗

which is why both systems ran in batches — parallel throughput converts a low per-attempt success probability into a high per-batch one. Two caveats matter in practice: proposals drawn from the same model on the same pool are correlated, so realised batch success rates fall short of this bound, and a larger batch is only worth its cost if throughput, not sample count, is the bottleneck actually being relieved.

Every negative result must be fed back, not discarded. In A-Lab, failed synthesis attempts updated the ARROWS3 route-prediction model for subsequent targets rather than being logged and ignored [5]. This is the step most manual, human-paced active-learning efforts skip under time pressure — running the next-best experiment is easy to prioritise over updating the model with the result of the last one, and skipping it converts an active-learning loop back into a static ranked list re-run periodically.

The loop needs a stopping rule stated in advance. Both systems ran for a fixed calendar duration, not until a target was hit, because in practice “run until the model is satisfied” is not a stopping rule the model can supply — expected improvement approaches zero as candidates are exhausted, but real budgets are constrained by instrument time and reagent cost, not by acquisition-function convergence. Decide the budget before starting: a fixed number of batches, a fixed number of days, or a fixed spend, so that a project without early results has a defined point at which the loop’s design, not just its luck, gets reviewed.

A robotic liquid-handling arm with its gripper closed around one sample vial, caught mid-lift out of a rack of model-proposed candidates toward an open slot on the synthesis carousel, a rejected vial from the previous round sitting apart in a tray marked for the model
Figure 1. An active-learning loop is not the model choosing; it is the model proposing, the robot executing, and the failed attempts going back to retrain the proposer.

The instructive part of both papers is how unglamorous the mechanism is once stated plainly: a ranking with an uncertainty term, a batch size chosen against a real bottleneck, a feedback path that is never allowed to be optional, and a budget decided before the first experiment runs.

What an AI medical device submission actually needs

Practitioners moving from a research model toward a device intended for clinical use in the United States are, in effect, moving from a modelling problem to an evidence-assembly problem, and the evidence categories are laid out in specific FDA documents rather than left to inference.

The starting reference is the joint set of ten guiding principles for Good Machine Learning Practice, issued in October 2021 by the FDA together with Health Canada and the UK’s Medicines and Healthcare products Regulatory Agency [9]. Read as a checklist rather than a philosophy statement, the principles that most directly shape what evidence a submission needs to contain are:

Independent training and test datasets. The guidance is explicit that training and test data must be independently sourced to the greatest extent possible — the same institutional, temporal and leakage concerns covered in the first section of this guide are, in the medical device context, a regulatory requirement rather than a best practice, and reviewers will ask how independence was ensured.

Representative clinical study populations and datasets. A submission needs to characterise, with evidence, whether the data used to train and validate the device reflects the population, disease severity range, and clinical sites in which it will actually be used. A model validated on a narrow demographic or a single acquisition protocol carries a burden to demonstrate — not simply assert — that its performance generalises to the intended use population.

Human-AI team performance, tested under clinically relevant conditions. For devices intended to support rather than replace a clinician’s judgement, the relevant endpoint is how accurately the human-plus-model system performs under real clinical workflow, not the model’s standalone accuracy on an idealised bench benchmark. This changes what a validation study needs to measure: a reader study with clinicians using the device as intended, against the range of inputs, artefacts and edge cases it will actually encounter.

Two more recent guidance documents matter specifically for how a device is allowed to change after clearance, which is a distinguishing feature of AI-enabled devices relative to traditional ones. The FDA’s final guidance on Predetermined Change Control Plans, issued in December 2024, lets a submission include a pre-specified, pre-authorised plan for future modifications — most often periodic retraining on new data — so that qualifying updates do not each require a new marketing submission [8]. A PCCP has three components a submission must supply concretely rather than in outline: a description of the specific planned modifications (what will change, and what will not), the methodology used to develop, validate and implement each modification (including the data sources and acceptance criteria), and an assessment of the impact each modification could have on the device’s benefit–risk profile. Writing a credible PCCP is, in practice, the same discipline as designing the leakage-resistant split and the active-learning stopping rule from the earlier sections: decide up front what evidence would justify a change, and what evidence would not.

A regulatory-affairs desk with a submission binder open, one tabbed divider labelled for a change control plan caught mid-insertion between already-filed sections, a benchtop device housing set to one side
Figure 2. A submission is not one dossier; it is several separate arguments — representative data, an independent test set, and a plan for what happens after clearance — assembled and cross-referenced.

Finally, FDA’s public list of authorised AI/ML-enabled medical devices is worth treating as primary source material, not background reading [7]. It records, device by device, which pathway was used — 510(k) clearance, De Novo authorisation, or premarket approval — and which predicate device a submission was compared against. Reading several submissions in a practitioner’s own device category before writing a new one remains the most efficient way to calibrate what “sufficient” evidence looks like, since the guidance describes categories of evidence without specifying exact study designs.

Building reproducibility into the pipeline, not bolting it on afterward

Reproducibility failures in computational science are rarely caused by dishonesty; they are caused by pipelines that were never designed to be re-run by anyone, including their own authors six months later. Two widely cited frameworks lay out what a pipeline needs to capture, and the discipline in both cases is to build the capture in from the first script rather than reconstruct it retrospectively when a reviewer asks.

Sandve, Nekrutenko, Taylor and Hovig’s ten rules remain the most concrete practitioner-level checklist available [10]: track the exact sequence of steps that produced a result automatically rather than by memory; never alter data by hand outside a recorded script; archive the exact version of every program used, not just its name; keep every custom script under version control; and — the rule most often skipped under deadline pressure — record random seeds wherever a stochastic process is involved, so a “close enough” reproduction is not the ceiling on what is achievable.

In practice this maps onto four pieces of infrastructure a computational science group should have before a headline result is generated, not after:

Environment capture. A pipeline’s software environment — language runtime, library versions, system dependencies — needs to be captured as an artefact (a container image, a lockfile, or an equivalent) alongside the code, not described in a methods paragraph. “We used scikit-learn” is not an environment; a pinned version list or a built container image is.

Data and code versioned together, with the pairing recorded. A specific model checkpoint or result table needs to be traceable to the exact commit of the code and the exact version or hash of the dataset that produced it. Code version control without a corresponding data version — or vice versa — leaves half of the provenance chain unrecoverable.

A workflow, not a folder of scripts run in an order somebody remembers. Explicit workflow definitions — whatever the specific tool — record the dependency graph between steps, so that “run script three before script five” is enforced by the pipeline rather than by institutional memory that leaves with whoever wrote it.

Findability and reuse of the underlying data, which is the specific concern the FAIR principles were written to address: that data be Findable, Accessible, Interoperable and Reusable by both humans and machines, with rich enough metadata that a dataset remains usable once its original creator is no longer available to explain it [11]. A pipeline that produces beautifully versioned code operating on an unlabelled, undocumented data file has only solved half the problem.

A pipeline-archive rack of dated drive caddies with one caddy caught mid-slide into its bay, a handwritten environment-snapshot tag only half attached, the previous quarter's caddies already seated and locked
Figure 3. Reproducibility is not a backup; it is code, data, environment and random seed locked together at the moment a result was produced, so the same run can be repeated on demand later.

The return on this discipline is not only defensive. A well-versioned pipeline is also what makes a PCCP’s “methodology for validating each modification,” discussed in the previous section, into something a reviewer can actually audit: the ability to point to exactly which code, data and environment produced last quarter’s validation numbers, and to rerun that exact combination on demand, is the same infrastructure whether the audience is a co-author, a regulator, or the researcher’s own future self debugging a discrepancy.

Reading a vendor’s claims like a reviewer, not a customer

Every problem covered so far — leakage-resistant splits, closed active-learning loops, evidence-grade validation, versioned reproducibility — is a checklist a practitioner can apply to a vendor’s product, not only to their own work, and doing so before adopting a tool is often the highest-leverage hour spent on a procurement decision.

The clearest cautionary case remains IBM’s Watson for Oncology. A 2018 investigation based on internal IBM documents found the system had, in multiple documented instances, produced “unsafe and incorrect” treatment recommendations, and traced the failure to its training: a small number of hypothetical or synthetic patient cases built by specialists at one institution, rather than real, diverse patient data validated against outcomes across the institutions where it was later deployed [12]. Every principle in the previous section — representative data, independent test conditions, realistic clinical evaluation — maps onto exactly what was missing. The lesson is not that the underlying idea was unsound; it is that a plausible-sounding system reached deployment without this guide’s evidence discipline applied at any stage, and its customers had no independent way to know that before real patients were affected.

Kapoor and colleagues’ REFORMS checklist, developed by consensus among nineteen researchers spanning computer science, statistics and biomedical fields, formalises the questions a reviewer — or a procurement evaluator — should be asking of any machine-learning-based scientific claim, organised as 32 questions across the full lifecycle of a study: problem specification, data quality, the modelling approach itself, the evaluation, and the conclusions actually supported by the results [2]. Adapted specifically to evaluating a vendor’s claims rather than reviewing a paper, the highest-value subset of that checklist is:

Ask specifically how the benchmark’s train/test split was constructed, using the vocabulary from the first section of this guide. “Cross-validated accuracy” on a benchmark split at random, on molecular or biological data, is close to meaningless without knowing whether the split respected scaffold, temporal or batch dependency; a vendor unable to answer the question precisely has likely not asked it either.

Ask whether the reported number came from a held-out set the vendor’s own team never touched during development, ideally curated or held by an independent third party. A benchmark a vendor curated, tuned against, and reports on itself is a demonstration, not independent evidence, no matter how large the reported number is.

Ask for prospective validation, not only retrospective backtesting. A model can be tuned, consciously or not, to a historical dataset’s idiosyncrasies. A prospective test — deploying the frozen model against data that does not yet exist at freeze time — is the same discipline recommended earlier for a practitioner’s own temporal splits, and a vendor should be able to describe one, even a small one.

A vendor-evaluation bench with an independent benchmark drive caught mid-plug into a test rig, the vendor's own demonstration drive disconnected and set aside on a walnut-wood tray next to its printed claims sheet
Figure 4. A claimed benchmark run on the vendor's own data is a demo, not evidence; the number that matters is the one measured on a held-out set the vendor never touched.

Ask what population and conditions the validation data represents, and compare that explicitly to the intended deployment population. A model validated at one hospital system, one instrument vendor, or one geographic population carries no automatic guarantee of transferring to another, and the burden of demonstrating transfer sits with whoever is proposing to deploy it there, not with the tool’s original benchmark.

Ask what the plan is for detecting and responding to performance drift after deployment, and whether that plan is written down anywhere resembling the PCCP structure described earlier, even for a research tool with no regulatory obligation to have one. A vendor with a considered answer to this question has usually thought seriously about the other four; a vendor without one usually has not.

None of these questions requires the evaluator to be a machine-learning specialist. They require the evaluator to insist on the same evidence a rigorous internal validation would need to produce, and to treat a vendor’s own benchmark the same way this guide has argued a practitioner should treat their own: as a claim to be checked, not a result to be trusted because it is stated confidently.

After deployment: the assumptions still need watching

Clearance, publication, or a successful pilot is a starting condition, not an ending one: every preceding section describes an assumption that can silently stop holding after deployment. The population a device sees in production can drift from its validation population. The space a model was actively learned against can shift as it moves into new chemistry or biology. A pipeline’s dependency versions can update in ways that change results without anyone editing a line of analysis code.

A monitoring desk with a dated log rack of past model-version cards, one new signed-off card caught mid-insertion into the current slot, an amber drift-alert light glowing quietly beside a live deployment terminal
Figure 5. Clearance is a starting condition, not an ending; the change-control plan filed with the submission is what says which drifts trigger a review and which do not.

The PCCP framework’s core insight — decide in advance which changes are expected and how they will be evaluated, rather than discovering the need for review only when something has visibly gone wrong — generalises well beyond regulated devices. A practical minimum for any deployed AI-for-science system: a dated log of model versions and the data each was validated against, an explicit and monitored signal for when the deployment population has drifted meaningfully from the validation population, and a standing decision about what magnitude of drift triggers a full re-validation rather than a shrug. Teams that build this before deployment treat it as routine bookkeeping; teams that build it only after a public failure treat it as incident response, at far higher cost and with far less credibility recovered per hour spent.

Predictions, and what would falsify them

These are forecasts, kept explicitly separate from the sourced analysis above. Horizon: August 2029.

One. Scaffold-, cluster-, or time-aware splitting will become the default expectation in computational chemistry and biology venues, with random splits requiring explicit justification rather than the reverse. Disconfirmed if leading venues in the field are still routinely accepting random-split validation as sufficient evidence for generalisation claims without comment from reviewers.

Two. Predetermined Change Control Plans, or an equivalent pre-authorised update mechanism, will extend beyond the FDA’s medical device pathway into other scientific-AI regulatory contexts, because the underlying problem — models meant to keep learning after initial validation — is not unique to medical devices. Disconfirmed if, by the horizon date, PCCP-equivalent mechanisms remain confined to the FDA medical device pathway with no analogous adoption elsewhere.

Three. Independent, third-party-held benchmark datasets for scientific-AI claims will become a distinguishing credibility signal that procurement processes explicitly request, in the way independent security audits became a standard vendor request over the previous decade. Disconfirmed if procurement and adoption decisions in this space continue to rely predominantly on vendor-reported, vendor-curated benchmark numbers with no material shift toward third-party or customer-run validation.

Four. The gap between published active-learning results and their reported hit or success rates will narrow as more groups publish negative results and failed synthesis attempts alongside successes, because the current literature is dominated by success-reporting systems like the two profiled in this guide. Disconfirmed if the ratio of published successful-loop papers to published failed-loop or negative-result papers in this area has not measurably shifted by the horizon date.

None of these four requires a capability breakthrough. They follow from a pattern already visible in the sources this guide relies on: the fields doing this work best have made splitting, closing the loop, assembling evidence, and versioning the pipeline into infrastructure rather than individual discipline.

What to take away

An AI-for-science result is only as strong as five ordinary decisions: the split that determined whether the held-out score means anything, the loop that determined whether predictions ever met a real measurement, the evidence assembled for whoever approves deployment, the versioning that determines whether a result can be reproduced on demand, and the scrutiny applied to any claim — including the practitioner’s own — before it is trusted. Each is checkable, each has a documented failure mode when skipped, and each has a documented fix. Treat all five as required infrastructure for the project, not as due diligence performed once the interesting modelling work is already finished.