Scientific institutions were not designed in one sitting. Peer review, the modern journal, the funded laboratory, and the registered protocol each accreted over centuries or decades through a mix of negotiated custom, statistical measurement, and, more recently, deliberate experiment. Ask why any one of them works the way it does today, and the answer depends heavily on which discipline you ask. A historian of science will hand you an archive. A metascientist will hand you a spreadsheet of ten thousand papers. A reform-minded editor will hand you the results of a randomized trial in which some manuscripts were reviewed before the data existed and some were not. All three are doing legitimate empirical work on the same subject. None of them is doing the other two’s job, and the differences are not a matter of taste — they are differences in what kind of causal claim the evidence can support at all.

This article lays those three approaches side by side: archival institutional history, quantitative metascience, and direct reform experimentation. It compares them on the dimensions that actually separate them — the strength of causal evidence each can produce, how far their conclusions generalize, and how directly their findings translate into a decision an editor, funder, or society council could act on tomorrow. The goal is not to declare a winner. Each approach answers a question the others cannot, and the strongest claims in this field are the ones the three happen to agree on independently.

Three ways of asking the same question

Consider one concrete institutional fact: the referee report. Today, a manuscript submitted to almost any serious journal is read by two or three anonymous specialists before an editor decides whether to publish it. Why does this exist, and did it make science better?

ADVERTISEMENT

An institutional historian answers by reconstructing sequence and negotiation. The Royal Society’s Philosophical Transactions, founded in 1665 under secretary Henry Oldenburg, ran for nearly a century under one editor’s personal judgment before a council committee took over review duties in the 1750s [4]. Formal, systematic peer review — two named Fellows producing signed written reports, a process proposed in 1831 by William Whewell in response to competition from French journals — did not arrive until the 1830s, nearly two hundred years after the journal’s founding [3]. That timeline is not incidental trivia; it tells you that “peer review” as practiced today is a nineteenth-century layer added onto a seventeenth-century institution for reasons specific to its moment — rising submission volume, professional rivalry, a need for external legitimacy that a single editor’s word could no longer supply. The evidence for this account is documentary: council minutes, correspondence, the text of Whewell’s proposal. It is a causal story built from a chain of specific decisions by named people, each decision traceable to an antecedent pressure.

An open bound society minute-book under a document-scanning cradle, a loupe resting on the page.
Figure 1. Archival institutional history reconstructs decisions—like the Royal Society's shift to signed referee reports—from the paper trail those decisions left behind.Image prompt and art direction by Brecht Corbeel; generation pending.

A metascientist answers a different question with the same subject: does peer review, as currently practiced across thousands of journals, correlate with better outcomes — fewer retractions, higher citation impact, less inflated significance? This requires no archive at all; it requires a large sample of papers, reviews, and outcomes, scored uniformly and compared statistically. Robert Merton’s mid-century sociology of science had already proposed that norms like organized skepticism functioned as a control system for the community as a whole, rather than as a property of any single review event [1] [2] — a claim that, unlike the Royal Society narrative, is stated at the level of the whole institution and is testable, in principle, against population data rather than a single documented episode.

A reform experimenter answers a third, narrower question: if you change one specific feature of review — say, requiring that the study design be evaluated before the results are known — does a measurable outcome change? This is the only one of the three approaches that assigns conditions rather than merely observing them, and it is correspondingly the only one that can claim experimental, rather than observational or archival, causal warrant.

Approach one: archival institutional history

Archival history’s evidentiary base is documents — minute books, correspondence, submission ledgers, published editorial policies — read in sequence to reconstruct why an institution took the shape it did. Steven Shapin and Simon Schaffer’s Leviathan and the Air-Pump is the field’s most influential demonstration of what this method can show: reading the Boyle–Hobbes dispute over the air pump as a contest over what would count as a properly witnessed experimental fact, they argued that the very category of “experiment” as a public, replicable, socially certified event was itself a seventeenth-century political and rhetorical achievement, not a timeless feature of nature-study [5]. That is a strong, specific, well-evidenced claim about one controversy in one decade — and it is exactly the kind of claim that only close reading of primary sources can support, because it depends on the particular arguments particular people made to particular audiences.

The strength of this approach is resolution: it can show mechanism — who argued what, against whom, and why a particular compromise stuck — in a way no aggregate statistic can. Its limitation is generalizability. A reconstruction of the Royal Society’s transition to signed referee reports, however well documented, is a claim about that institution in that century. Whether the same causal logic (rising volume, competition, a need for external legitimacy) explains the analogous transition in German university science, or in twentieth-century American funding panels, is a separate historical question requiring its own archive. Archival history rarely produces a number you can compare across a hundred institutions, and it is not built to. What it produces instead is irreplaceable texture: the reasons a decision was defensible to the people who made it, which quantitative approaches typically cannot recover at all, because the record they work from has already stripped the deliberation out and kept only the outcome.

ADVERTISEMENT
A rack-mounted server and terminal displaying a citation network graph still partly rendered.
Figure 2. Metascience treats institutions as populations: citation counts, retraction records, and replication attempts scored across thousands of cases at once.Image prompt and art direction by Brecht Corbeel; generation pending.

Approach two: quantitative metascience

Metascience treats the scientific literature itself as a dataset. Derek de Solla Price’s 1963 Little Science, Big Science is usually credited as the founding text of this approach: rather than studying any single controversy, Price applied statistical methods to the aggregate growth of publications, citations, and authorship, treating “science” as a measurable population with its own regularities [6]. That move — counting rather than reading — is what makes bibliometrics, citation analysis, and later retraction-and-replication tracking a genuinely distinct approach rather than history with more footnotes.

Two results illustrate both the power and the limits of this style of evidence. John Ioannidis’s 2005 argument, “Why Most Published Research Findings Are False,” used a formal probabilistic model of study power, effect size, and researcher degrees of freedom to argue that, under realistic assumptions about how many hypotheses are tested and how flexibly, a majority of published positive findings in many fields are expected to be false even when every individual study was conducted honestly [7]. That is a claim about aggregate expected reliability across a research field — not a claim about any specific paper, and not something an archival case study of one journal could have produced, because the argument is inherently about a whole population of tested hypotheses.

The 2015 Open Science Collaboration project supplied the empirical counterpart: it attempted direct replications of one hundred studies published in three major psychology journals and found that, while 97 percent of the original studies had reported statistically significant results, only 36 percent of the replications did, with replicated effect sizes about half the original magnitude on average [8]. This is metascience at its strongest — a pre-registered, large-sample, uniformly scored measurement of an institutional outcome (how often published significant results recur) across an entire discipline’s flagship venues at once.

The limitation is construct validity: a citation count is a proxy for influence, not influence itself; a “successful replication” is defined by a statistical threshold that itself embeds assumptions the metascience literature continues to argue over. And the causal story stays thin. The Open Science Collaboration’s 36 percent figure tells you a great deal about the rate of a phenomenon and very little, on its own, about which specific institutional feature — sample sizes, incentive structures, undisclosed flexibility in analysis, publication bias, or something else entirely — produced that rate in any one case. Metascience is extremely good at showing that something is wrong at scale and comparatively weak at showing, on its own, which lever fixes it. That is the third approach’s job.

A printed tray of retraction-database rows beside the bibliometric terminal, one row still printing.
Figure 3. Large retraction and replication databases let metascientists ask population-level questions no single archive or trial could answer.Image prompt and art direction by Brecht Corbeel; generation pending.

Approach three: direct reform experiments

Where archival history observes and metascience measures, the reform-experiment tradition intervenes. The clearest example is the registered report. Christopher Chambers introduced the format at the journal Cortex in 2013: authors submit their study design and analysis plan for peer review before collecting data, and the journal commits to publishing the result — significant, null, or messy — based on the quality of the design rather than the attractiveness of the outcome [9]. This directly targets one specific mechanism metascience had flagged as a likely contributor to the replication problem: publication bias against null results, which pressures researchers toward flexible analysis in search of something significant enough to print. Reviews of the accumulating registered-report literature report that outcome-blind review is associated with a markedly higher share of null results reaching publication compared with the standard literature, and with fewer of the specific analytic red flags — selective outcome reporting, post hoc hypothesis framing — that Ioannidis’s model implicated [10]. By the mid-2020s the format had spread to several hundred journals across disciplines, itself a metascience-trackable fact about the reform’s own adoption [9].

A stapled preregistration protocol pinned under a paperweight beside a sealed envelope marked for post-registration results.
Figure 4. Registered-report trials manipulate one institutional lever directly—review before results exist—and measure whether the outcome changes.Image prompt and art direction by Brecht Corbeel; generation pending.

This is the closest thing science-of-science research has to a controlled experiment: one variable (timing of review relative to results) manipulated, while the surrounding institutional context — same journals, same broad research communities, same era — stays comparatively constant. That gives registered-report evaluations the strongest internal causal warrant of the three approaches: when the outcome differs between standard and registered tracks within the same journal and the same period, the timing-of-review variable is a far more defensible candidate explanation than in either archival or purely observational metascience work.

ADVERTISEMENT

The same design that buys internal validity costs external validity. A trial run within psychology and neuroscience journals that adopted registered reports voluntarily says less than it might seem to about fields, journals, or countries that did not adopt the format — the adopters are self-selected, likely already somewhat reform-minded, which is exactly the kind of confound a randomized trial is supposed to remove but a voluntary editorial adoption cannot. Funding-side reform experiments face an even sharper version of this problem: an analysis of NIH’s Center for Scientific Review peer-review network found that reviewer scores correlate with later measured impact overall, but that the relationship all but disappears within the top-scoring tier of proposals, where most of the actual funding decisions are made [11] — a finding that limits how far any single funding-reform trial’s results can be expected to travel to a different funding tier, a different disease area, or a different national system. Direct experiments answer “does this specific lever, pulled this specific way, in this specific setting, move this specific outcome” with unusual confidence — and answer almost nothing else with the same confidence.

Two identical lab-tray apparatus set up side by side for a matched replication run, one dust cover half-lifted.
Figure 5. A direct replication attempt only tests whether one result recurs; it cannot by itself explain why an institution produced the result in the first place.Image prompt and art direction by Brecht Corbeel; generation pending.

Setting the three side by side

The honest comparison is not which approach is right but which question each is built to answer, and what it costs to get that answer.

Causal evidence strength runs, roughly, from weakest to strongest as: single-case archival narrative, cross-sectional metascience correlation, and controlled reform trial. But that ordering reverses almost exactly for external validity and mechanism-level richness. The archival account of the Royal Society’s 1830s shift to signed reports explains, in granular and well-evidenced detail, why that specific institution made that specific change — a claim a randomized trial could never produce, because no trial randomizes seventeenth-century learned societies. The Open Science Collaboration’s 36 percent figure generalizes across a discipline’s flagship journals in a way no single registered-report trial or single archival case can, precisely because it is a population measurement rather than an intervention or a narrative. The registered-report trials generalize the least of the three in raw scope — they cover the journals and fields that adopted the format — but within that scope they license the strongest specific causal claim: this procedural change, this outcome shift.

Actionability follows a similar split. An editor deciding whether to adopt registered reports at their own journal has a direct, mechanism-matched precedent to act on [9] [10]. A funding agency asking whether to redesign study-section scoring has a metascience literature suggesting the current system’s discriminating power is weak exactly where the money is actually allocated [11] [12], which argues for experimentation but does not by itself specify which alternative scoring scheme would do better — that gap is precisely where a controlled funding-reform trial would need to sit, and few of the scale needed have yet been run. A historian’s account of why the Royal Society changed its review process in 1831 offers no direct prescription for a modern funding panel at all; it offers something different and still valuable — a demonstration that review procedures are contingent, negotiated, and changeable, which is itself a useful corrective to the assumption that today’s peer review is a natural or inevitable feature of science rather than one nineteenth-century institution’s solution to its own nineteenth-century problem.

Where the three approaches agree

The strongest claims in this literature are not the ones any single approach makes alone but the places where archival, quantitative, and experimental evidence converge from independent directions. All three traditions agree that peer review, in its familiar signed-or-anonymous-referee form, is a comparatively recent addition bolted onto a much older publication system rather than a feature present from the scientific revolution onward — the archival record dates the shift to the 1830s [3], and the sheer diversity of review formats metascience surveys internationally today, from single-blind to open to registered, is hard to reconcile with the idea of one settled, ancient procedure. All three also converge, from different evidentiary bases, on the claim that publication bias against null or messy results has been a real and consequential feature of the system: Merton’s mid-century description of organized skepticism as a communal ideal already implied that individual incentives and communal norms could diverge [1]; Ioannidis’s model formalized how selective reporting degrades aggregate reliability [7]; and the registered-report trials show that removing the bias mechanically, by reviewing before outcomes exist, shifts the published record’s composition in the predicted direction [10]. When a historical account, a statistical model, and a controlled trial independently point at the same mechanism, that convergence is stronger evidence than any one of the three would be alone — not because the methods stopped disagreeing about everything else, but because they disagree about enough else that agreement on this one point is unlikely to be a shared artifact of any single method’s blind spot.

What this comparison does not settle

None of the three approaches, singly or combined, settles the normative question of how much reform any given institution should undertake, or how fast. A funding agency told that top-tier scores barely discriminate future impact [11] still has to decide what to do with that fact — abandon fine-grained scoring, run its own randomized allocation trial, fund a broader portfolio, or conclude the measurement problem lies elsewhere — and metascience evidence alone under-determines that choice as sharply as archival evidence under-determines whether the Royal Society’s 1831 reform was, on balance, good for science rather than merely different from what preceded it. Predictions about where institutional reform goes next — for instance, whether registered reports become the default rather than the exception across the sciences within the next decade — are genuinely uncertain: the observable indicator to watch is the format’s adoption rate outside the early-adopter journals and fields that chose it voluntarily, and the clearest disconfirming signal would be adoption plateauing well below a majority of major journals in a field for a sustained period, which would suggest the format solves a problem specific to certain research communities rather than a general one. Where the three approaches genuinely disagree — for example, on how much of the replication shortfall metascience attributes to publication bias versus other causes such as underpowered designs or field-specific effect-size inflation — that disagreement should be reported as a live methodological dispute among researchers who read the same underlying evidence differently, not resolved by asserting one camp’s number as settled fact.