A Piece of Software Has Now Run Google’s Own Infrastructure Longer Than It Has Waited for a Referee
Google’s own account of AlphaEvolve says the quiet part in one clause easy to read past. Announced on the DeepMind blog on May 14, 2025, the post describes a scheduling heuristic the system evolved for Borg, the software that allocates computing jobs across Google’s data centers, as a solution “now in production for over a year,” one that “continuously recovers, on average, 0.7% of Google’s worldwide compute resources” [11]. Read that plainly: before the company said a public word about how the heuristic was found, it had already been reallocating a measurable slice of one of the world’s largest compute fleets for more than twelve months, unannounced. Now read the calendar forward from the announcement instead of back from it. As of this writing, close to sixteen months after that blog post, the technical report describing AlphaEvolve’s method has not been accepted at any peer-reviewed venue [10] — no journal listing, no conference proceedings, nothing beyond the arXiv preprint itself. A system already old enough to have quietly earned Google over a year of production credit has now gone almost as long again waiting, still unchecked by any outside referee, for someone to formally examine the paper describing how it works.
The document a referee would actually need to read did not even exist yet when the announcement ran. AlphaEvolve’s technical report was not posted to arXiv until June 16, 2025, more than a month after the DeepMind blog post first described the system publicly [10]. So the honest floor on AlphaEvolve’s own review clock is not the blog post’s date but the report’s — closer to fifteen months than sixteen — and even that more forgiving count changes nothing about the destination: as of this writing, no journal or conference has accepted it.
AlphaEvolve is the extreme case, not an outlier manufactured for effect. It is one entry in a six-system pattern running from February 2023 to the present, across the papers and blog posts that built the trajectory now widely described as an AI-fueled revolution in computational evolution: Language Model Crossover, Promptbreeder, Eureka, FunSearch, AlphaEvolve, and The AI Scientist. Of the six, exactly one had cleared peer review at the moment it first reached an audience. The other five did not — and two of them, as of this writing, still have not, at any point since. This article is not an argument that any of the six results is wrong. It is a ledger, built entirely from each paper’s own dated metadata — arXiv submission history, conference proceedings pages, journal issue records, repository timestamps — of when, if ever, a referee actually looked at the work the field had already started citing.
Six Papers, One Question the Field Answered Before Anyone Checked
The naive assumption is easy to state and easy to hold without ever testing it: a result famous enough to reshape a research trajectory must have passed peer review, because that is what peer review is for. Nothing about the six systems above requires that assumption to be true, and nothing about how quickly each one got cited or built upon waited to find out. Language Model Crossover proposed and named the mechanism — an LLM turning a handful of examples into a new candidate via few-shot prompting — that Promptbreeder, Eureka, FunSearch, and AlphaEvolve would each go on to use under their own branding, testing it across five separate representations in its own experiments: binary strings, sentences, equations, image-generation prompts, and Python code [1]. Promptbreeder demonstrated that mechanism could evolve its own mutation instructions, not just the prompts those instructions mutate [3]. Eureka pointed the same pattern at reward-function design for robotics and reported outperforming human experts on 83 percent of tested tasks with a 52 percent average improvement [5]. FunSearch pointed it at open combinatorial mathematics. AlphaEvolve turned it into a general-purpose coding agent running inside Google’s own production stack. The AI Scientist wired the pattern through an entire research pipeline, idea to review, and produced a full paper for under fifteen dollars [12]. Every one of these facts is already established, cited, and built upon elsewhere. None of that citation and adoption activity paused to ask whether a referee had signed off first.
The test this piece runs is narrow and entirely mechanical: for each system, find the date it first reached a public audience — an arXiv posting or a company blog announcement, whichever came first — and find the date, if any, it was accepted at a formal peer-reviewed venue: a journal issue, a conference’s own proceedings listing. The gap between those two dates is the legitimacy lag. A system that never clears that gap, at all, as of this writing, is marked as such rather than assigned an artificial number. This is bookkeeping, not a claim about scientific quality — a paper’s peer-review status says nothing directly about whether its results replicate — but it is bookkeeping the field’s own behavior has made worth doing, because the field visibly did not wait for the answer before treating each of these six systems as established fact.
Lay the Six Dates Side by Side and the Split Falls Exactly at Twelve Months
Lay the six dates side by side and the pattern is not subtle.
| System | First public date | First peer-reviewed date | Lag |
|---|---|---|---|
| FunSearch | Dec 14, 2023 (DeepMind blog) [9] | Dec 14, 2023 (Nature) [8] | 0 |
| Eureka | Oct 19, 2023 (arXiv) [5] | May 7, 2024 (ICLR 2024) [6] | ~6.5 months |
| Promptbreeder | Sep 28, 2023 (arXiv) [3] | Jul 21, 2024 (ICML 2024) [4] | ~9.75 months |
| Language Model Crossover | Feb 23, 2023 (arXiv) [1] | Nov 28, 2024 (ACM TELO) [2] | ~21 months |
| AlphaEvolve | May 14, 2025 (DeepMind blog) [11] | none as of this writing | ~15.75 months, open |
| The AI Scientist | Aug 12, 2024 (arXiv) [12] | none as of this writing | ~25 months, open |
“First public date” means whichever came first between a preprint’s own arXiv submission timestamp and a company’s own blog announcement, never a news outlet’s coverage date or a social-media mention — the point is to measure the paper’s own dated footprint, not how loudly the field talked about it. “First peer-reviewed date” means acceptance at a named, indexed journal issue or conference proceedings, confirmed through the publisher’s or organizer’s own record rather than an author’s claim of an upcoming venue. Every date in that table is a primary-source fact, not an estimate: FunSearch’s blog post and its companion Nature paper share the same publication date, December 14, 2023, confirmed independently through both DeepMind’s own announcement and Nature’s indexed record [9] [8] — the rarest condition on this list, a result whose first public appearance and first peer-reviewed appearance are the same event. That same-day alignment is not an accident of timing; it is what happens when a lab chooses to submit its own announcement to a journal’s review process before saying anything publicly at all, then times the blog post to the journal’s own publication date rather than the other way around. FunSearch is the one system on this ledger where the referee’s opinion was a precondition for the announcement rather than an afterthought to it — the sequence every one of the other five reversed. Eureka’s arXiv posting on October 19, 2023, and its ICLR 2024 acceptance, with the conference itself running May 7 through 11, 2024, are both drawn from the paper’s own listing and the conference’s official dates [5] [6]. Promptbreeder’s arXiv posting on September 28, 2023, and its ICML 2024 appearance — published in the conference’s own proceedings at pages 13481 through 13544, with the conference itself running July 21 through 27, 2024 — come from the same kind of direct record [3] [4]. Language Model Crossover took the longest road of the three that eventually arrived: posted February 23, 2023, revised twice, and not formally published until the ACM Transactions on Evolutionary Learning and Optimization issued it online on November 28, 2024 — an interval of just over twenty-one months, confirmed through the publisher’s own indexed record [1] [2].
The arithmetic that matters most sits in what the three completed cases have in common and where the other three stop. Among the systems that did eventually clear a formal review — FunSearch, Eureka, and Promptbreeder — the lag runs 0, roughly 6.5, and roughly 9.75 months, a median of about six and a half months. That is a real, if uneven, pace: unglamorous, occasionally slow, but not indefinite. The other three tell a different story. Language Model Crossover exceeded a full year before it was reviewed at all. AlphaEvolve, at close to sixteen months and still counting, has not. The AI Scientist, at roughly twenty-five months and still counting, has not either. Split the six exactly where the twelve-month line falls and the result is precise rather than approximate: three of six systems were formally reviewed within a year of first reaching an audience; three were not, and as of this writing two of those three still have no review at all. Half, exactly, on the record as it stands today.
What Stood In for a Referee, Case by Case
A missing referee is not automatically a missing check. Each of the six systems accumulated some form of scrutiny during its own lag, and the kind of scrutiny differs enough from case to case that flattening them into one undifferentiated “unreviewed” bucket would understate what actually happened.
FunSearch needed no substitute; its formal review and its public debut were the same event, an unusually clean case among the six. Eureka’s stand-in during its six-and-a-half-month wait was public and inspectable rather than institutional: the authors released the full training and evaluation code on GitHub the same month as the arXiv posting, under an MIT license, with the repository’s own description crediting the paper’s central 83-percent and 52-percent figures [7] — anyone with access to the same simulated robotics environments could, in principle, rerun the comparison against human-engineered rewards without waiting for ICLR’s verdict. Promptbreeder’s nine-and-three-quarter-month wait had no equivalent open-code substitute as visible from its own arXiv listing, and its case is closer to the ordinary condition of conference-cycle latency: a paper that eventually cleared review through the normal channel, just slower than FunSearch’s same-day case and faster than Language Model Crossover’s twenty-one months. That ordinary latency is worth naming precisely, because it recurs below in a very different context: a wait of that length for a major-conference decision is not, on its own, evidence of anything unusual in how this field operates. It is close to the length of a single standard conference review cycle, the same kind of interval that separates any large-lab language-model paper’s arXiv posting from its own eventual proceedings appearance, evolutionary computation or not.
AlphaEvolve’s stand-in is the most unusual of the six, precisely because it inverts the normal order of operations. Ordinarily, a claim earns scrutiny before it earns deployment. AlphaEvolve’s Borg-scheduling result had already been running against real infrastructure, moving real jobs, recovering a real and continuously measured share of compute, for over a year before the public claim was even made [11]. That is a genuine form of validation — a result that fails in production tends to get noticed by the people running the production system, referee or no referee — but it is validation by an interested party checking its own tool against its own infrastructure, not an outside check against the paper’s broader claims about mathematics and algorithm design, which is where AlphaEvolve’s other reported results, evolved code for problems well outside Google’s own scheduler, still have no comparable production-scale test standing in for review.
The AI Scientist’s stand-in is the most formal of the five non-simultaneous cases, and it is worth describing precisely rather than treating as generic “some people looked into it.”
The Independent Check That Never Touched This Office’s Own Stamp
Six months after The AI Scientist’s original arXiv posting, Joeran Beel, Min-Yen Kan, and Moritz Baumgart ran the full pipeline again, on a domain the original paper never touched, and posted their findings to arXiv on February 20, 2025 [13]. Their evaluation is not a peer-review report in the formal sense — nobody asked them to referee the original submission, and their study happened well after The AI Scientist had already made its impact — but it is the closest thing this ledger has to an outside audit conducted by researchers with no stake in the original claim, and its own numbers are concrete rather than impressionistic. Running idea generation through seven complete manuscripts cost $42 total, about $6 each, cheaper than the original paper’s own $15 figure [13] — a point in the original claim’s favor. Of twelve proposed experiments, five failed outright on coding errors, a 42 percent failure rate on the original paper’s own disclosed weak point, now measured independently on unfamiliar ground [13]. On the review stage specifically, the finding is unambiguous: “the reviewer agent did not identify any of the truly serious issues that we highlighted,” and the evaluators’ summary judgment is blunter still — “the AI Scientist cannot critically assess its own results. It fails to detect methodological flaws or logical inconsistencies, making it unsuitable for autonomous scientific inquiry” [13]. Their characterization of the manuscripts themselves does not soften the point: “the quality of its manuscripts currently aligns with that of an unmotivated undergraduate student rushing to meet a deadline” [13] — a verdict the evaluators reached only after a real cost of their own time: five hours setting the system up, then fifteen implementing an experimental template and a further ten brainstorming and reviewing output, which — excluding that initial setup — they total as “25 hours of human effort… approximately 3.5 hours per manuscript” [13].
Here is the detail that belongs in this specific ledger and nowhere else in this series: that evaluation itself did not stay outside the review system it was auditing. It was later published under an ACM digital-object identifier in SIGIR Forum in June 2025 [13] — the checking work carried out on a different desk, with a different instrument, eventually passed through a stamp of its own, even though the system it examined still has not. The original paper that made The AI Scientist famous remains, as of this writing, exactly where it started: an arXiv preprint with no journal reference and no conference listing, twenty-five months and counting.
This Is Not a Peculiarity of Evolutionary Computation
The obvious objection to everything above is that it proves nothing specific to this trajectory: fast-moving machine learning research runs on preprints generally, and singling out six evolutionary-computation-adjacent systems for an unusually long wait would be unfair if every other corner of contemporary AI research waits just as long, or longer. That objection deserves a real test, not a shrug, so here is one: four of the most-cited large language model papers of the same general era, chosen for fame rather than for convenience, run through the identical two-date measurement.
GPT-3’s paper, “Language Models are Few-Shot Learners,” posted to arXiv on May 28, 2020, was accepted at NeurIPS 2020, which ran December 6 through 12 that year — a lag of about six months, comfortably inside the twelve-month line, and the fastest of the four [14]. PaLM tells a slower story with unusually precise documentation: posted to arXiv on April 5, 2022, the paper’s own published header records it as “Submitted 10/22; Revised 6/23; Published 8/23” in the Journal of Machine Learning Research — meaning Google did not even submit the paper to a journal for six months after posting it publicly, and the review process that followed ran another ten months past that, for a total lag of roughly sixteen months from first public appearance to print [16]. GPT-4’s technical report, posted March 15, 2023, has never been accepted at any peer-reviewed venue, as of this writing — the same absence found in AlphaEvolve’s and The AI Scientist’s rows above, now running past forty months [15]. LLaMA, posted February 27, 2023, sits in the same unreviewed category, at a comparable length of time [17].
Run the same twelve-month cutoff against this four-paper sample and one of four cleared it: GPT-3, at six months. PaLM, GPT-4, and LLaMA did not — a 75 percent miss rate, worse than the 50 percent found across the six systems this article has been auditing. This is a small, hand-picked sample chosen for name recognition rather than random selection, and it should not be read as a rigorous field-wide baseline; a systematic survey across hundreds of contemporaneous preprints could move either number.
One plausible institutional explanation for the pattern, offered here as a reasonable account rather than a tested finding, is that a frontier lab’s own incentive to submit a paper for formal review is weakest exactly when the result is already doing its intended work without one — GPT-4 launched as a paid commercial product the same week its technical report first appeared, with the report itself withholding architecture and training details on stated competitive grounds [15], and AlphaEvolve’s scheduling heuristic was already recovering measurable compute inside Google’s own fleet for over a year before the company wrote a public paper about it at all. A journal’s endorsement adds credibility a lab already has by other means; it does not add revenue, deployment, or a fixed publication deadline the way a product launch or a conference’s own submission cycle does. That account would predict exactly the split visible above — venue-bound results like Eureka and Promptbreeder, tied to a conference’s own clock, clear review at a normal pace, while results that already function as shipped infrastructure or a shipped product have the weakest institutional reason to bother. But within that limit, the honest reading is not that the evolutionary-computation-and-language-model trajectory has an unusually broken relationship with peer review. It is that the entire frontier of large language model research currently treats formal review as optional, slow when it happens at all, and this trajectory’s own three-of-six record is, if anything, a slightly better showing than four of the field’s most famous papers manage among themselves. The residue that survives this comparison is narrower than the opening pages of this article might have suggested, and more useful for being narrower: half of six systems missing a referee for more than a year is a real, checkable condition of how this specific trajectory reached its audience, but it is not evidence that this trajectory is unusually informal. It is evidence that informal is now the frontier’s default setting, and this trajectory is simply not exempt from it.
What a Blog Post Should Buy a Reader’s Trust, and What It Shouldn’t
None of this licenses treating an unreviewed claim as equivalent to a checked one, and none of it licenses the opposite mistake of treating peer review’s absence as proof of a hidden flaw. What the ledger actually supports is a specific, transferable habit for reading the next company blog post in this trajectory, rather than a verdict on any one paper’s underlying correctness. A number that has cleared a real referee — FunSearch’s cap-set and bin-packing results, Eureka’s and Promptbreeder’s eventual ICLR and ICML acceptances — carries a form of institutional backing the other three do not yet have, whatever their other merits. A number that has not, however impressive the demonstration attached to it, is a claim currently vouched for only by the party making it and, in AlphaEvolve’s case specifically, by that same party’s own production infrastructure — a real form of stress-testing, but not an independent one. The AI Scientist’s case adds a further, sharper lesson: an outside check can exist, be rigorous, be quotable, and even clear a formal venue of its own, while the original claim it examined remains permanently unreviewed by the standard everyone assumed it had already met.
The falsifiable version of this piece’s own claim is simple to state and cheap to re-run at any later date: if AlphaEvolve’s technical report or The AI Scientist’s original paper clears a peer-reviewed venue before this ledger is next updated, that row moves from open to closed, and the fraction changes from three-of-six to four-of-six or better. Until one of those two rows closes, the honest number to carry away from this trajectory is not “the field doesn’t check its work.” It is narrower and more precise: half of the six papers that built this record took more than a year to face a referee, two of those three still haven’t, and the field spent that entire interval, in every case, citing, building on, and shipping the results anyway.