Institutions are built piecewise, not designed
Popular accounts of science describe its institutions as though they arrived whole: journals that always sent submissions out for expert review, laboratories that always trained graduate students on university payrolls, publishing that was always locked behind subscriptions until the internet fixed it. The documented history is different in a specific, useful way — each institution was assembled under local pressure, piece by piece, and the pieces that survived were the ones that solved a recurring problem rather than the ones anyone designed on paper. This briefing walks through three such mechanisms: manuscript refereeing, the research laboratory, and open-access publishing, and reads each against what metascience research now says about how well the resulting system actually verifies claims.
Mechanism one: how peer review actually hardened
Fact, from the historical record. Philosophical Transactions, launched in March 1665 by the Royal Society’s first secretary Henry Oldenburg, is the longest continuously running scientific journal [3]. Its early issues were not refereed submissions in any modern sense — they were edited excerpts of Oldenburg’s own correspondence with natural philosophers across Europe, book reviews, and accounts of observations, compiled and printed under his personal editorial judgment [3].
Formal refereeing did not exist yet as a named practice. Historians Noah Moxham and Aileen Fyfe, working from the Royal Society’s own archives, trace how the “function of refereeing” emerged gradually from the social practices of arranging a gentlemen’s learned society’s meetings and publications during the late eighteenth and nineteenth centuries — not as a single reform but as an accumulation of ad hoc habits: asking a knowledgeable member to look a paper over, recording an opinion, sometimes overruling it [1]. A documented turning point came in 1752, when a crisis in the Society’s affairs prompted an institutional takeover of the journal; a 21-person committee began collectively selecting papers, replacing Oldenburg-style individual editorial discretion with committee review [1, 3]. Fyfe and colleagues, examining the Society’s referee reports and correspondence through 1965, find that even after the 1830s brought more systematic and rigorous expert review, the volume of manuscripts needing review, and the administrative machinery for managing it, kept expanding well into the twentieth century — peer review as a mass, routinised, forms-and-deadlines bureaucratic process is substantially a twentieth-century phenomenon layered on top of a nineteenth-century foundation [2].
Analysis. The lesson is not that peer review is fake or recent in the dismissive sense sometimes claimed online. It is that peer review was never one invention with one founding date — it is a stack of successive fixes (an editor’s personal judgment, then committee selection, then named external referees, then formal report forms, then, in the twentieth century, blind review and structured criteria) each added to patch a failure mode the previous layer did not catch. That layered history matters for reading current debates about review quality: complaints that today’s peer review is inconsistent, slow, or gameable are complaints about the latest layer of a centuries-long patching process, not evidence that a once-perfect system has decayed.
Mechanism two: how the research laboratory became an institutional form
Fact. Before the 1820s, advanced chemical training in much of Europe resembled apprenticeship in a master’s private workshop: one senior chemist, a small number of assistants, informally arranged. Justus von Liebig, appointed extraordinary professor of chemistry at the University of Giessen in 1824 after study with Joseph Louis Gay-Lussac in Paris, built something different: a large, systematic teaching laboratory in which many students at once received hands-on training in quantitative chemical analysis as a structured curriculum rather than an individual apprenticeship [4, 5]. He secured government funding from the Darmstadt authorities specifically to build the laboratory, and by the time he left Giessen for Munich in 1852, more than 700 students of chemistry and pharmacy had trained there [5].
Why this counts as an institutional discontinuity rather than just a bigger workshop. The Giessen laboratory decoupled two things that apprenticeship had bundled together: a master’s personal reputation, and a training capacity that could scale. Once Giessen demonstrated that a laboratory could reliably produce competent analytical chemists in quantity, other universities across Europe and later the United States copied the model directly — hiring their own professors, building their own benches, replicating the same standardized bench-scale training pipeline [4]. Historians of science have described this as arguably the single most important element in the international rise of graduate research education across academic fields generally, not chemistry alone [5].
Analysis, separated from claim. It is a documented fact that Giessen’s laboratory model was widely imitated. It is an analytical judgment, not a settled fact with a single causal proof, that this imitation caused the broader rise of research-university graduate education rather than merely coinciding with and accelerating trends — state investment in applied chemistry, industrial demand for trained analysts — already under way. Both readings are consistent with the documented imitation pattern; the strong causal-primacy claim should be treated as a widely held historical interpretation rather than a demonstrated mechanism.
Mechanism three: how open access actually changed the pipeline
Fact. In 1991, Paul Ginsparg, then at Los Alamos National Laboratory, automated an existing informal email distribution list of string-theory preprints — begun in 1989 by Joanne Cohn — into what became the Los Alamos e-print archive, later arXiv.org [6]. Ginsparg has written that he spent a few afternoons that summer writing the original software; the server was, for years, physically housed under his own desk before moving with him to Cornell in 2001 [6]. By its twentieth anniversary, arXiv had become the default first-circulation point for large swaths of physics, mathematics, and computer science, hosting drafts before, sometimes instead of, and occasionally without ever going through, formal journal peer review [7].
What this actually changed, mechanically. Before preprint servers and open-access journals, the sequence was fixed: finish the work, submit to one journal, wait through review (often many months), publish, then the result becomes visible to the field. arXiv (and the broader open-access movement it anticipated) decoupled circulation from formal review — a result becomes visible to the field on a timescale of hours to days, while formal peer review, if it happens at all before or after, runs on its own separate and usually much longer clock. This is a genuine institutional discontinuity: it is not merely “faster publishing,” it is two previously fused steps — public circulation and formal vetting — becoming separable and running in parallel or in either order.
Why that discontinuity has a cost, per current metascience. Decoupling circulation from vetting does not remove the need for vetting; it just relocates the moment when it happens, and sometimes removes the moment where it happens at all for a given piece of work. The Open Science Collaboration’s 2015 replication project is not about publishing speed, but it measures the downstream consequence of any system, fast or slow, that lets unverified claims accumulate faster than they can be checked: attempting high-powered direct replications of 100 psychology studies originally published across three journals, the project found that while 97% of the original studies reported a statistically significant effect, only 36% of the replications did, and replicated effect sizes were on average about half the magnitude of the originals [8]. That result predates arXiv-style preprinting’s spread into psychology, but it establishes the baseline problem that any faster-circulation system inherits and can worsen if verification capacity does not scale with it: publication, in any pipeline, is not the same event as a claim being checked.
What the three mechanisms share
None of these three institutions — refereeing, the teaching laboratory, open-access circulation — was planned as a complete system from a blueprint. Each was a local fix for a specific bottleneck (Oldenburg’s correspondence outgrowing one editor’s attention; Gay-Lussac-style apprenticeship not scaling to demand; journal review cycles outpacing the speed at which a field wanted to see new work) that then got copied elsewhere because it worked well enough, not because it was optimal. Reading current complaints about any of the three — review is slow, laboratories are underfunded, preprints let bad work circulate unchecked — as evidence of decay misreads the history. These are, and always were, patched systems; the relevant scientific and editorial question in each period is the same one Fyfe’s group asks of the Royal Society’s own management records: whether the verification capacity of the moment is scaling with the volume of claims moving through it [2]. That question has no permanent answer — it has to be re-asked of every generation’s version of the pipeline, including the current one built on preprint servers and rapid online circulation.
A conditional note, clearly marked as such. If open-access and preprint systems continue expanding circulation faster than review and replication capacity scale, current metascience findings such as the Reproducibility Project’s replication rate suggest the checked-to-total-claims ratio could keep declining across more fields over the next decade — not a prediction of collapse, but a scenario contingent on review and replication infrastructure not receiving comparable investment to publishing infrastructure. The observable indicator to watch is whether replication and registered-report capacity (funded slots, journal formats, institutional credit for replication work) grows as a share of total published output; the disconfirming condition would be a stable or rising replication rate in fields with expanding preprint use, which would indicate institutions adapted verification capacity alongside circulation capacity rather than falling behind it.