Then in LinkedIn: Write article → click into the body → paste (Ctrl+V). Headings, links and images come with it. The title usually pastes as the first line — cut it into LinkedIn's title field. back to the article

AlphaEvolve Ran Inside Google for a Year Before Any Referee Saw a Word of It

Only one of the six papers that built the AI-fueled merger between evolutionary computation and large language models had passed peer review the moment it first reached an audience.

A wall of six numbered pigeonhole mail slots in an archival office, each holding one envelope, only one envelope bearing a fresh red "received" ink stamp

Six papers reached their first public audience between February 2023 and May 2025. Only one of the six envelopes in this office was already stamped by a referee when it arrived. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Abstract

Six papers and blog posts built the AI-fueled merger between evolutionary computation and large language models between February 2023 and May 2025: Language Model Crossover, Promptbreeder, Eureka, FunSearch, AlphaEvolve and The AI Scientist. Only one, FunSearch, had cleared peer review at the moment it first reached a public audience — the DeepMind blog post and the Nature paper landed the same day. This article builds a legitimacy-lag ledger from each system's own arXiv, conference, journal and repository metadata: three of the six eventually passed formal review, at a median gap of about six and a half months; the other three took more than twenty months, or, in AlphaEvolve's and The AI Scientist's cases, still have not as of this writing. It tests that finding against four of the most-cited large language model papers of the same era — GPT-3, PaLM, GPT-4 and LLaMA — to check whether the gap is a peculiarity of this specific research trajectory or the ordinary condition of frontier AI publishing, and states plainly what a reader should do with a company blog post's claim while a referee's opinion remains, in some cases years later, unavailable.

A Piece of Software Has Now Run Google’s Own Infrastructure Longer Than It Has Waited for a Referee

Google’s own account of AlphaEvolve says the quiet part in one clause easy to read past. Announced on the DeepMind blog on May 14, 2025, the post describes a scheduling heuristic the system evolved for Borg, the software that allocates computing jobs across Google’s data centers, as a solution “now in production for over a year,” one that “continuously recovers, on average, 0.7% of Google’s worldwide compute resources” [11]. Read that plainly: before the company said a public word about how the heuristic was found, it had already been reallocating a measurable slice of one of the world’s largest compute fleets for more than twelve months, unannounced. Now read the calendar forward from the announcement instead of back from it. As of this writing, close to sixteen months after that blog post, the technical report describing AlphaEvolve’s method has not been accepted at any peer-reviewed venue [10] — no journal listing, no conference proceedings, nothing beyond the arXiv preprint itself. A system already old enough to have quietly earned Google over a year of production credit has now gone almost as long again waiting, still unchecked by any outside referee, for someone to formally examine the paper describing how it works.

The document a referee would actually need to read did not even exist yet when the announcement ran. AlphaEvolve’s technical report was not posted to arXiv until June 16, 2025, more than a month after the DeepMind blog post first described the system publicly [10]. So the honest floor on AlphaEvolve’s own review clock is not the blog post’s date but the report’s — closer to fifteen months than sixteen — and even that more forgiving count changes nothing about the destination: as of this writing, no journal or conference has accepted it.

AlphaEvolve is the extreme case, not an outlier manufactured for effect. It is one entry in a six-system pattern running from February 2023 to the present, across the papers and blog posts that built the trajectory now widely described as an AI-fueled revolution in computational evolution: Language Model Crossover, Promptbreeder, Eureka, FunSearch, AlphaEvolve, and The AI Scientist. Of the six, exactly one had cleared peer review at the moment it first reached an audience. The other five did not — and two of them, as of this writing, still have not, at any point since. This article is not an argument that any of the six results is wrong. It is a ledger, built entirely from each paper’s own dated metadata — arXiv submission history, conference proceedings pages, journal issue records, repository timestamps — of when, if ever, a referee actually looked at the work the field had already started citing.

Six Papers, One Question the Field Answered Before Anyone Checked

The naive assumption is easy to state and easy to hold without ever testing it: a result famous enough to reshape a research trajectory must have passed peer review, because that is what peer review is for. Nothing about the six systems above requires that assumption to be true, and nothing about how quickly each one got cited or built upon waited to find out. Language Model Crossover proposed and named the mechanism — an LLM turning a handful of examples into a new candidate via few-shot prompting — that Promptbreeder, Eureka, FunSearch, and AlphaEvolve would each go on to use under their own branding, testing it across five separate representations in its own experiments: binary strings, sentences, equations, image-generation prompts, and Python code [1]. Promptbreeder demonstrated that mechanism could evolve its own mutation instructions, not just the prompts those instructions mutate [3]. Eureka pointed the same pattern at reward-function design for robotics and reported outperforming human experts on 83 percent of tested tasks with a 52 percent average improvement [5]. FunSearch pointed it at open combinatorial mathematics. AlphaEvolve turned it into a general-purpose coding agent running inside Google’s own production stack. The AI Scientist wired the pattern through an entire research pipeline, idea to review, and produced a full paper for under fifteen dollars [12]. Every one of these facts is already established, cited, and built upon elsewhere. None of that citation and adoption activity paused to ask whether a referee had signed off first.

The test this piece runs is narrow and entirely mechanical: for each system, find the date it first reached a public audience — an arXiv posting or a company blog announcement, whichever came first — and find the date, if any, it was accepted at a formal peer-reviewed venue: a journal issue, a conference’s own proceedings listing. The gap between those two dates is the legitimacy lag. A system that never clears that gap, at all, as of this writing, is marked as such rather than assigned an artificial number. This is bookkeeping, not a claim about scientific quality — a paper’s peer-review status says nothing directly about whether its results replicate — but it is bookkeeping the field’s own behavior has made worth doing, because the field visibly did not wait for the answer before treating each of these six systems as established fact.

Lay the Six Dates Side by Side and the Split Falls Exactly at Twelve Months

Lay the six dates side by side and the pattern is not subtle.

“First public date” means whichever came first between a preprint’s own arXiv submission timestamp and a company’s own blog announcement, never a news outlet’s coverage date or a social-media mention — the point is to measure the paper’s own dated footprint, not how loudly the field talked about it. “First peer-reviewed date” means acceptance at a named, indexed journal issue or conference proceedings, confirmed through the publisher’s or organizer’s own record rather than an author’s claim of an upcoming venue. Every date in that table is a primary-source fact, not an estimate: FunSearch’s blog post and its companion Nature paper share the same publication date, December 14, 2023, confirmed independently through both DeepMind’s own announcement and Nature’s indexed record [9] [8] — the rarest condition on this list, a result whose first public appearance and first peer-reviewed appearance are the same event. That same-day alignment is not an accident of timing; it is what happens when a lab chooses to submit its own announcement to a journal’s review process before saying anything publicly at all, then times the blog post to the journal’s own publication date rather than the other way around. FunSearch is the one system on this ledger where the referee’s opinion was a precondition for the announcement rather than an afterthought to it — the sequence every one of the other five reversed. Eureka’s arXiv posting on October 19, 2023, and its ICLR 2024 acceptance, with the conference itself running May 7 through 11, 2024, are both drawn from the paper’s own listing and the conference’s official dates [5] [6]. Promptbreeder’s arXiv posting on September 28, 2023, and its ICML 2024 appearance — published in the conference’s own proceedings at pages 13481 through 13544, with the conference itself running July 21 through 27, 2024 — come from the same kind of direct record [3] [4]. Language Model Crossover took the longest road of the three that eventually arrived: posted February 23, 2023, revised twice, and not formally published until the ACM Transactions on Evolutionary Learning and Optimization issued it online on November 28, 2024 — an interval of just over twenty-one months, confirmed through the publisher’s own indexed record [1] [2].

A hand-bound ledger book open on a sloped desk showing a column of dated entries in fountain-pen ink, the bottom entry's ink still wet and darker than the dried rows above it

Figure 1. Six rows. One closed the same day it opened. The rest of this ledger is built from nothing but each paper's own dated metadata, entered here in the order the record actually allows. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.

The arithmetic that matters most sits in what the three completed cases have in common and where the other three stop. Among the systems that did eventually clear a formal review — FunSearch, Eureka, and Promptbreeder — the lag runs 0, roughly 6.5, and roughly 9.75 months, a median of about six and a half months. That is a real, if uneven, pace: unglamorous, occasionally slow, but not indefinite. The other three tell a different story. Language Model Crossover exceeded a full year before it was reviewed at all. AlphaEvolve, at close to sixteen months and still counting, has not. The AI Scientist, at roughly twenty-five months and still counting, has not either. Split the six exactly where the twelve-month line falls and the result is precise rather than approximate: three of six systems were formally reviewed within a year of first reaching an audience; three were not, and as of this writing two of those three still have no review at all. Half, exactly, on the record as it stands today.

What Stood In for a Referee, Case by Case

A missing referee is not automatically a missing check. Each of the six systems accumulated some form of scrutiny during its own lag, and the kind of scrutiny differs enough from case to case that flattening them into one undifferentiated “unreviewed” bucket would understate what actually happened.

FunSearch needed no substitute; its formal review and its public debut were the same event, an unusually clean case among the six. Eureka’s stand-in during its six-and-a-half-month wait was public and inspectable rather than institutional: the authors released the full training and evaluation code on GitHub the same month as the arXiv posting, under an MIT license, with the repository’s own description crediting the paper’s central 83-percent and 52-percent figures [7] — anyone with access to the same simulated robotics environments could, in principle, rerun the comparison against human-engineered rewards without waiting for ICLR’s verdict. Promptbreeder’s nine-and-three-quarter-month wait had no equivalent open-code substitute as visible from its own arXiv listing, and its case is closer to the ordinary condition of conference-cycle latency: a paper that eventually cleared review through the normal channel, just slower than FunSearch’s same-day case and faster than Language Model Crossover’s twenty-one months. That ordinary latency is worth naming precisely, because it recurs below in a very different context: a wait of that length for a major-conference decision is not, on its own, evidence of anything unusual in how this field operates. It is close to the length of a single standard conference review cycle, the same kind of interval that separates any large-lab language-model paper’s arXiv posting from its own eventual proceedings appearance, evolutionary computation or not.

AlphaEvolve’s stand-in is the most unusual of the six, precisely because it inverts the normal order of operations. Ordinarily, a claim earns scrutiny before it earns deployment. AlphaEvolve’s Borg-scheduling result had already been running against real infrastructure, moving real jobs, recovering a real and continuously measured share of compute, for over a year before the public claim was even made [11]. That is a genuine form of validation — a result that fails in production tends to get noticed by the people running the production system, referee or no referee — but it is validation by an interested party checking its own tool against its own infrastructure, not an outside check against the paper’s broader claims about mathematics and algorithm design, which is where AlphaEvolve’s other reported results, evolved code for problems well outside Google’s own scheduler, still have no comparable production-scale test standing in for review.

A wall calendar pad with a small pile of torn-off leaves beneath it and the currently displayed leaf caught half-torn, still attached along one edge

Figure 2. Twenty-one months for one paper. Six and a half for another. Sixteen and still climbing for a third. Same pad; different rates of tearing. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.

A "received" rubber stamp resting dry on its ink pad beside a separate envelope whose exposed edge has visibly yellowed, no ink mark anywhere on the envelope itself

Figure 3. Sixteen months for one of these two papers, twenty-five for the other, and this stamp still has not touched either envelope. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.

The AI Scientist’s stand-in is the most formal of the five non-simultaneous cases, and it is worth describing precisely rather than treating as generic “some people looked into it.”

The Independent Check That Never Touched This Office’s Own Stamp

Six months after The AI Scientist’s original arXiv posting, Joeran Beel, Min-Yen Kan, and Moritz Baumgart ran the full pipeline again, on a domain the original paper never touched, and posted their findings to arXiv on February 20, 2025 [13]. Their evaluation is not a peer-review report in the formal sense — nobody asked them to referee the original submission, and their study happened well after The AI Scientist had already made its impact — but it is the closest thing this ledger has to an outside audit conducted by researchers with no stake in the original claim, and its own numbers are concrete rather than impressionistic. Running idea generation through seven complete manuscripts cost $42 total, about $6 each, cheaper than the original paper’s own $15 figure [13] — a point in the original claim’s favor. Of twelve proposed experiments, five failed outright on coding errors, a 42 percent failure rate on the original paper’s own disclosed weak point, now measured independently on unfamiliar ground [13]. On the review stage specifically, the finding is unambiguous: “the reviewer agent did not identify any of the truly serious issues that we highlighted,” and the evaluators’ summary judgment is blunter still — “the AI Scientist cannot critically assess its own results. It fails to detect methodological flaws or logical inconsistencies, making it unsuitable for autonomous scientific inquiry” [13]. Their characterization of the manuscripts themselves does not soften the point: “the quality of its manuscripts currently aligns with that of an unmotivated undergraduate student rushing to meet a deadline” [13] — a verdict the evaluators reached only after a real cost of their own time: five hours setting the system up, then fifteen implementing an experimental template and a further ten brainstorming and reviewing output, which — excluding that initial setup — they total as “25 hours of human effort… approximately 3.5 hours per manuscript” [13].

A plainer second desk off to one side holding a magnifying loupe tilted mid-lift over a marked-up printout covered in handwritten annotations, no stamp or ledger present on this desk

Figure 4. The independent check that found this system's review stage missing every flaw it was shown never passed through this office's own stamp at all. The checking still happened, on a different desk, under a different instrument. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Here is the detail that belongs in this specific ledger and nowhere else in this series: that evaluation itself did not stay outside the review system it was auditing. It was later published under an ACM digital-object identifier in SIGIR Forum in June 2025 [13] — the checking work carried out on a different desk, with a different instrument, eventually passed through a stamp of its own, even though the system it examined still has not. The original paper that made The AI Scientist famous remains, as of this writing, exactly where it started: an arXiv preprint with no journal reference and no conference listing, twenty-five months and counting.

This Is Not a Peculiarity of Evolutionary Computation

The obvious objection to everything above is that it proves nothing specific to this trajectory: fast-moving machine learning research runs on preprints generally, and singling out six evolutionary-computation-adjacent systems for an unusually long wait would be unfair if every other corner of contemporary AI research waits just as long, or longer. That objection deserves a real test, not a shrug, so here is one: four of the most-cited large language model papers of the same general era, chosen for fame rather than for convenience, run through the identical two-date measurement.

GPT-3’s paper, “Language Models are Few-Shot Learners,” posted to arXiv on May 28, 2020, was accepted at NeurIPS 2020, which ran December 6 through 12 that year — a lag of about six months, comfortably inside the twelve-month line, and the fastest of the four [14]. PaLM tells a slower story with unusually precise documentation: posted to arXiv on April 5, 2022, the paper’s own published header records it as “Submitted 10/22; Revised 6/23; Published 8/23” in the Journal of Machine Learning Research — meaning Google did not even submit the paper to a journal for six months after posting it publicly, and the review process that followed ran another ten months past that, for a total lag of roughly sixteen months from first public appearance to print [16]. GPT-4’s technical report, posted March 15, 2023, has never been accepted at any peer-reviewed venue, as of this writing — the same absence found in AlphaEvolve’s and The AI Scientist’s rows above, now running past forty months [15]. LLaMA, posted February 27, 2023, sits in the same unreviewed category, at a comparable length of time [17].

Run the same twelve-month cutoff against this four-paper sample and one of four cleared it: GPT-3, at six months. PaLM, GPT-4, and LLaMA did not — a 75 percent miss rate, worse than the 50 percent found across the six systems this article has been auditing. This is a small, hand-picked sample chosen for name recognition rather than random selection, and it should not be read as a rigorous field-wide baseline; a systematic survey across hundreds of contemporaneous preprints could move either number.

One plausible institutional explanation for the pattern, offered here as a reasonable account rather than a tested finding, is that a frontier lab’s own incentive to submit a paper for formal review is weakest exactly when the result is already doing its intended work without one — GPT-4 launched as a paid commercial product the same week its technical report first appeared, with the report itself withholding architecture and training details on stated competitive grounds [15], and AlphaEvolve’s scheduling heuristic was already recovering measurable compute inside Google’s own fleet for over a year before the company wrote a public paper about it at all. A journal’s endorsement adds credibility a lab already has by other means; it does not add revenue, deployment, or a fixed publication deadline the way a product launch or a conference’s own submission cycle does. That account would predict exactly the split visible above — venue-bound results like Eureka and Promptbreeder, tied to a conference’s own clock, clear review at a normal pace, while results that already function as shipped infrastructure or a shipped product have the weakest institutional reason to bother. But within that limit, the honest reading is not that the evolutionary-computation-and-language-model trajectory has an unusually broken relationship with peer review. It is that the entire frontier of large language model research currently treats formal review as optional, slow when it happens at all, and this trajectory’s own three-of-six record is, if anything, a slightly better showing than four of the field’s most famous papers manage among themselves. The residue that survives this comparison is narrower than the opening pages of this article might have suggested, and more useful for being narrower: half of six systems missing a referee for more than a year is a real, checkable condition of how this specific trajectory reached its audience, but it is not evidence that this trajectory is unusually informal. It is evidence that informal is now the frontier’s default setting, and this trajectory is simply not exempt from it.

What a Blog Post Should Buy a Reader’s Trust, and What It Shouldn’t

None of this licenses treating an unreviewed claim as equivalent to a checked one, and none of it licenses the opposite mistake of treating peer review’s absence as proof of a hidden flaw. What the ledger actually supports is a specific, transferable habit for reading the next company blog post in this trajectory, rather than a verdict on any one paper’s underlying correctness. A number that has cleared a real referee — FunSearch’s cap-set and bin-packing results, Eureka’s and Promptbreeder’s eventual ICLR and ICML acceptances — carries a form of institutional backing the other three do not yet have, whatever their other merits. A number that has not, however impressive the demonstration attached to it, is a claim currently vouched for only by the party making it and, in AlphaEvolve’s case specifically, by that same party’s own production infrastructure — a real form of stress-testing, but not an independent one. The AI Scientist’s case adds a further, sharper lesson: an outside check can exist, be rigorous, be quotable, and even clear a formal venue of its own, while the original claim it examined remains permanently unreviewed by the standard everyone assumed it had already met.

The falsifiable version of this piece’s own claim is simple to state and cheap to re-run at any later date: if AlphaEvolve’s technical report or The AI Scientist’s original paper clears a peer-reviewed venue before this ledger is next updated, that row moves from open to closed, and the fraction changes from three-of-six to four-of-six or better. Until one of those two rows closes, the honest number to carry away from this trajectory is not “the field doesn’t check its work.” It is narrower and more precise: half of the six papers that built this record took more than a year to face a referee, two of those three still haven’t, and the field spent that entire interval, in every case, citing, building on, and shipping the results anyway.

Sources

  1. Elliot Meyerson, Mark J. Nelson, Herbie Bradley, Adam Gaier, Arash Moradi, Amy K. Hoover, and Joel Lehman. Language Model Crossover: Variation through Few-Shot Prompting. arXiv (2023). DOI: 10.48550/arXiv.2302.12170.
  2. Elliot Meyerson, Mark J. Nelson, Herbie Bradley, Adam Gaier, Arash Moradi, Amy K. Hoover, and Joel Lehman. Language Model Crossover: Variation through Few-Shot Prompting. ACM Transactions on Evolutionary Learning and Optimization (2024). DOI: 10.1145/3694791.
  3. Chrisantha Fernando, Dylan Banarse, Henryk Michalewski, Simon Osindero, and Tim Rocktäschel. Promptbreeder: Self-Referential Self-Improvement via Prompt Evolution. arXiv (2023). DOI: 10.48550/arXiv.2309.16797.
  4. Chrisantha Fernando, Dylan Banarse, Henryk Michalewski, Simon Osindero, and Tim Rocktäschel. Promptbreeder: Self-Referential Self-Improvement via Prompt Evolution. Proceedings of the 41st International Conference on Machine Learning (PMLR vol. 235) (2024).
  5. Yecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang, Osbert Bastani, Dinesh Jayaraman, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Eureka: Human-Level Reward Design via Coding Large Language Models. arXiv (2023). DOI: 10.48550/arXiv.2310.12931.
  6. Yecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang, Osbert Bastani, Dinesh Jayaraman, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Eureka: Human-Level Reward Design via Coding Large Language Models. International Conference on Learning Representations (ICLR 2024) (2024).
  7. Ma, Liang, Wang, Huang, Bastani, Jayaraman, Zhu, Fan, and Anandkumar. Eureka (official repository). GitHub (2023).
  8. Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M. Pawan Kumar, Emilien Dupont, Francisco J. R. Ruiz, Jordan S. Ellenberg, Pengming Wang, Omar Fawzi, Pushmeet Kohli, and Alhussein Fawzi. Mathematical discoveries from program search with large language models. Nature (2023). DOI: 10.1038/s41586-023-06924-6.
  9. Google DeepMind. FunSearch: Making New Discoveries in Mathematical Sciences Using Large Language Models. Google DeepMind (blog) (2023).
  10. Alexander Novikov, Ngân Vũ, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco J. R. Ruiz, Abbas Mehrabian, M. Pawan Kumar, Abigail See, Swarat Chaudhuri, George Holland, Alex Davies, Sebastian Nowozin, Pushmeet Kohli, and Matej Balog. AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv (2025). DOI: 10.48550/arXiv.2506.13131.
  11. Google DeepMind. AlphaEvolve: A Gemini-Powered Coding Agent for Designing Advanced Algorithms. Google DeepMind (blog) (2025).
  12. Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv (2024). DOI: 10.48550/arXiv.2408.06292.
  13. Joeran Beel, Min-Yen Kan, and Moritz Baumgart. Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?. ACM SIGIR Forum (2025). DOI: 10.1145/3769733.3769747.
  14. Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, and others. Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems 33 (NeurIPS 2020) (2020).
  15. OpenAI. GPT-4 Technical Report. arXiv (2023). DOI: 10.48550/arXiv.2303.08774.
  16. Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, and others. PaLM: Scaling Language Modeling with Pathways. Journal of Machine Learning Research (2023).
  17. Hugo Touvron, Thibaut Lavril, Gautier Izacard, and others. LLaMA: Open and Efficient Foundation Language Models. arXiv (2023). DOI: 10.48550/arXiv.2302.13971.

Originally published at https://absolutedigitalpublishers.com/articles/alphaevolve-ran-a-year-in-production-before-any-referee-saw-it.