Ask an economic historian what caused the productivity gains of the First or Second Industrial Revolution and the honest answer depends on which desk they sit at. A national-accounts researcher will pull up a growth-accounting decomposition and quote a share of output growth attributable to steam or electrification. A business historian will pull a specific firm’s ledgers and describe exactly how one mill manager rearranged shifts around a new engine. A comparative development economist will lay a dozen countries’ adoption dates side by side and ask why some took decades longer than others to get the same machine running at scale. All three are studying the same historical fact. None of the three methods can fully answer the other two’s question, and that is not a flaw to be fixed — it is what each method is built to do and not do.
This article compares three major approaches to studying industrial revolutions and general-purpose technologies (GPTs): quantitative cliometric growth accounting, firm-level business history built from archival records, and comparative cross-country development economics. It works through the same substantive terrain each approach covers — mechanization, the factory system, transport, electrification, productivity, skill demand, firm strategy, diffusion, and inequality — and asks, on each, what kind of evidence each tradition uses, what causal claim it can support, and where it runs out of road. It does not pick a winner. The three are complementary because they answer different questions with different evidence; ranking them against one another is a category error the same way ranking a microscope against a telescope would be.
What a “general-purpose technology” claim actually asserts
The concept itself comes from a specific methodological move. Timothy Bresnahan and Manuel Trajtenberg proposed that a small number of technologies — the steam engine, electric motor, and semiconductor among them — drive whole eras of growth because they combine three properties: pervasiveness across sectors, continued technical improvement over a long period, and “innovational complementarities,” meaning the technology’s payoff rises as the sectors that use it invest in complementary changes of their own [1]. That framing already tells you what kind of claim is being made: a GPT claim is a claim about the shape of a technology’s diffusion and its interaction with downstream reorganization, not simply a claim about how good the artifact is at its narrow original task.
This distinction matters because it is exactly where the three approaches split. Cliometric growth accounting is built to measure the aggregate output effect of that diffusion once it has happened. Business history is built to show what “complementary reorganization” actually consisted of, firm by firm. Comparative development economics is built to ask why the diffusion took the shape and pace it did in one place and not another. A GPT claim invokes all three mechanisms at once, so no one approach, alone, can fully test it.
Approach one: cliometric growth accounting
Cliometrics names the application of formal economic theory and statistical method to historical data — quantifying variables that earlier historical writing treated qualitatively, and then testing hypotheses against the resulting series [9]. Applied to a GPT question, the classic move is growth accounting: build a long time series of output, capital, labor, and (where possible) a measure of the specific technology’s capital stock or usage, then decompose measured growth into the share attributable to that input versus total factor productivity.
Nicholas Crafts’s growth-accounting study of steam power is the clean illustration of both the method’s power and its discipline. Using national accounts and estimates of steam’s diffusion into the British economy, Crafts found that steam’s contribution to growth was small before 1830 and reached its peak around a century after Watt’s key patents — only once high-pressure steam engines after 1850 delivered the efficiency and power-density needed to matter at national scale [2]. That is a genuinely falsifiable, quantified claim: it says a specific number, for a specific period, computed from a specific accounting identity, and it directly overturned an earlier, less quantified intuition that steam’s importance tracked its invention date rather than its later diffusion and refinement.
Paul David generalized the pattern into an explicit historical analogy. Comparing electrification’s slow, multi-decade productivity payoff to the contemporary (1990) computer productivity paradox, David argued that a GPT’s national productivity effect lags its invention for structural reasons — factories had to be rebuilt around unit-drive electric motors rather than simply swap a steam engine for a dynamo, and that rebuilding took a generation of capital turnover and organizational relearning [3]. This is cliometrics doing what it does best: turning a historical pattern into a comparable, quantifiable claim that then transfers as a testable hypothesis to a different era’s technology.
What this approach can support: precise, comparable estimates of aggregate contribution over time, testable against alternative accounting assumptions; direct falsification of claims that overstate an early technology’s importance; transferable structural hypotheses (adoption lag, complementary capital) that can be checked against other GPT episodes.
Where it runs out of road: growth accounting requires that the inputs be measurable at a national or sectoral level over long runs, which means it says very little about mechanism — why a particular mill manager waited fifteen years to convert to unit-drive motors is invisible to a series that only records aggregate capital stock. It also depends heavily on the quality and coverage of historical national accounts, which for most of the world before the twentieth century simply do not exist at the resolution the method wants, which is precisely the gap the next two approaches are built to fill from different directions.
Approach two: firm-level business history
Where cliometrics starts from national aggregates and works down, business history starts from a single firm’s surviving paper trail — minute books, wage ledgers, correspondence, patent files, engineering notebooks — and works outward to characterize how technology adoption actually happened inside an organization.
Alessandro Nuvolari’s reconstruction of steam-engine development in the British collieries is a case study in what this buys you. Rather than treat “steam engine efficiency” as a single number improving smoothly over time, archival work on the Cornish mining districts documents a specific institutional mechanism — engineers publicly reported engine performance data (through what is known in the literature as Lean’s Engine Reporter) in a competitive, semi-cooperative arrangement that let rival colliery engineers compare and improve designs without formal patent protection, a pattern historians have since generalized as “collective invention” [6]. No national accounts series could recover that mechanism; it is legible only from the archival record of who reported what to whom, and when.
This is also where firm-level evidence complicates a tidy cliometric story about labor markets. Robert Allen’s wage-and-output reconstruction for the British industrial revolution period documents “Engels’ pause” — a stretch from roughly 1790 to 1840 in which real wages for ordinary workers stagnated even as per-capita output rose sharply, with the gap accruing disproportionately to profits and rents [4]. That is itself a national-accounts (cliometric) finding, but explaining why the pause happened at the level of specific labor contracts, piece rates, and factory discipline requires exactly the kind of ledger-level evidence business history specializes in — how a specific mill set wages, enforced hours, and substituted machine-minders for skilled artisans task by task, which is the mechanism Daron Acemoglu and Pascual Restrepo’s later “task-based” framework for automation formalizes: technology can simultaneously displace workers from old tasks and create new ones, and which effect dominates depends on firm-level choices about task design, not only on the aggregate capital stock [7].
What this approach can support: causal mechanism at the level where decisions were actually made — specific adoption timing, specific organizational adaptation, specific bargaining over wages and tasks; discovery of institutional arrangements (like collective invention) invisible to any aggregate series; correction of overly smooth national narratives by showing the lumpy, contested, locally contingent reality underneath them.
Where it runs out of road: a single firm’s archive, however rich, is not a random sample of the economy, and business historians are candid that a well-documented firm may be well-documented precisely because it was unusual — larger, more literate in its record-keeping, or more successful than its neighbors, all of which bias any inference toward atypical cases. Generalizing from one colliery’s or one mill’s ledgers to a claim about “the” industrial revolution is an inferential leap the archival evidence itself cannot license; that step needs either many such case studies assembled comparably, or a return to the aggregate series the cliometric approach specializes in.
Approach three: comparative cross-country development economics
The third approach neither aggregates one country’s history nor drills into one firm’s ledgers; it lines up the same technology’s diffusion across many countries and asks what explains the differences in timing, speed, and downstream effect.
Diego Comin and Bart Hobijn’s CHAT dataset is the infrastructural example: a panel covering the historical adoption of over a hundred technologies across more than 150 countries since 1800, built specifically to let researchers test diffusion theories against comparable cross-country adoption-lag data rather than against any single country’s narrative [5]. The kind of question this supports is different in kind from either of the first two approaches: not “how much did steam add to British output” and not “how did this one firm adopt steam,” but “why did steam, or electrification, or any other GPT, reach country B twenty years after it reached country A, and does that lag track measurable differences in human capital, institutions, or market size?”
Comparative long-run GDP reconstruction supplies the other half of this approach’s evidence base. Stephen Broadberry, Johann Custodis, and Bishnupriya Gupta’s national-accounts-based comparison of British and Indian per-capita output from 1600 to 1871 found that India’s relative position — over 60 percent of British per-capita GDP around 1600 — had fallen to under 15 percent by 1871, and that the decline was already well underway in the seventeenth and eighteenth centuries, before Britain’s most intensive mechanization [8]. That finding reframes a GPT-diffusion question: it suggests that access to a specific general-purpose technology cannot by itself be the whole explanation for the divergence in outcomes, since a meaningful part of the gap in underlying capacity predates the technology’s arrival — a claim that only becomes visible by holding two countries’ long-run series up against each other on comparable terms.
What this approach can support: identification of which country-level variables correlate with adoption speed and downstream divergence across a large number of independent cases, giving it more generalizing power than a single firm study and more resolution on cross-country variation than an aggregate single-country accounting exercise; tests of whether a technology’s effects are conditional on prior institutional or human-capital endowments rather than automatic.
Where it runs out of road: cross-country panels built from historical statistics inherit every measurement problem in the underlying national accounts, multiplied across however many countries are in the panel, and the countries with the worst historical record-keeping are often exactly the ones whose divergence the researcher most wants to explain. Correlation across countries is also a weak tool for causal identification on its own — a lag correlated with, say, literacy rates does not establish that literacy caused the lag rather than both tracking some third factor, a limitation the comparative literature manages by combining panel evidence with detailed institutional case knowledge rather than resolving it away.
Where the three approaches actually check each other
The most useful way to see the complementarity is to watch what happens when all three are run on the same episode.
On mechanization and the factory system, cliometric growth accounting can tell you that manufacturing productivity rose and roughly by how much; it cannot tell you that the mechanism was a specific reorganization of the workday around continuous machine operation rather than, say, a shift in raw-material quality. Firm ledgers supply that mechanism directly, at the cost of not knowing whether the firm under study was representative. Comparative country data can then check whether the mechanism the firm evidence identifies (say, factory discipline enforced through fines and clocked hours) is present in other countries that mechanized at different speeds, which is a test neither of the other two approaches can run alone.
On transport and electrification, the same layering applies. Aggregate growth accounting is exactly the tool David used to notice that electrification’s productivity payoff lagged the technology by decades [3]; that lag itself becomes a fact to be explained, and explaining it took business-history-style evidence about how long it actually took to redesign a factory floor around unit-drive motors instead of a central shaft-and-belt system, plus comparative evidence on which countries’ capital-replacement cycles let them rebuild fastest.
On inequality, Allen’s wage-stagnation finding is itself a national-accounts (cliometric) result, but assessing whether “Engels’ pause” reflects a general property of early GPT diffusion or a specific feature of British labor institutions requires exactly the comparative check a cross-country panel is built to run, and exactly the firm-level detail on how wages and tasks were actually set that only ledgers preserve [4, 7].
Separating fact, disagreement, and open method question
A few things in this comparison are established findings, not points of dispute: that steam’s British growth contribution was concentrated after 1850 rather than at its invention [2]; that British real wages stagnated for several decades while output rose in the early industrial period [4]; that India’s per-capita output relative to Britain’s had already fallen substantially before Britain’s most intensive nineteenth-century mechanization [8].
Where genuine disagreement exists, it is largely about weighting rather than about the facts above: how much of a GPT’s aggregate effect should be attributed to the core technology itself versus the complementary reorganization it enables is a question growth accounting and business history answer with structurally different evidence, and reasonable researchers assign different shares to each depending on which archive or series they trust more at the margin. This is a live disagreement about interpretation, not a factual dispute, and it is not resolved by picking one method as more rigorous than the other — the methods are measuring different things.
An open method question, distinct from either of the above, is how far cross-country panel data like CHAT can be pushed toward causal claims about why adoption lagged, given that the historical statistics underlying such panels vary enormously in quality and completeness across countries and periods [5]. That is an active area of methodological development, not a settled fact, and treating a panel correlation as if it were a firm-level causal mechanism would overstate what the evidence supports.
A scenario, stated as a scenario
One plausible near-term direction, offered explicitly as scenario rather than forecast, is that large-scale digitization of firm archives (of the kind the planetary-scanner and archival-copy-stand work in this article’s own image world represents) could let researchers build panel datasets out of firm-level records at a scale that starts to blur the line between business history and comparative economics — testing mechanism-level hypotheses across hundreds of firms rather than one or two. The assumptions underneath that scenario are that enough firm archives survive in usable condition, that digitization and record-linkage costs keep falling, and that firms across different countries kept comparable enough categories of record for the linkage to be meaningful. The horizon for a first useful such panel, on current digitization trends discussed in the archival and cliometric literature, is plausibly the next one to two decades rather than the next few years. The observable indicator to watch is whether cliometric journals begin publishing studies with firm-level panels spanning dozens of firms across multiple countries rather than the current pattern of single-country aggregate series or single-firm case studies. The condition that would disconfirm the scenario is if firm-level record survival and comparability turn out to be too uneven across countries for meaningful linkage — in which case the three approaches would likely remain as methodologically separate as they are now, each still checking the others’ blind spots rather than merging into one.
A note on skills, and why the three approaches read “skill” differently
Skill demand is worth pausing on separately because it is the dimension where the three approaches most visibly talk past each other if a reader is not careful about which is which. A growth-accounting series can show that the wage premium for a broad occupational category — “skilled” versus “unskilled” — moved in a particular direction over a period, but the category itself is an aggregation choice made by whoever built the series, and it can obscure as much as it reveals; a rising average skill premium is consistent with several very different underlying stories. Firm-level ledgers show the underlying story directly for one workplace: which specific tasks a specific machine took over, which specific job titles disappeared from the payroll, and which new job titles appeared to operate, maintain, or supervise the machine. Acemoglu and Restrepo’s task-based framework was built precisely to formalize that firm-level pattern — automation of existing tasks displaces the workers who did them, while the same technology can simultaneously create new tasks that reinstate labor demand elsewhere in the same production process, and which of the two effects dominates in the aggregate series depends on the balance of task displacement against task creation that only firm-level or task-level evidence can actually show [7]. A comparative cross-country lens adds a third layer again: whether the same technology reinstates labor through new tasks at all may depend on a country’s existing human-capital base and institutions, which is exactly the kind of conditioning variable comparative panels are built to test and single-firm or single-country studies cannot.
None of this is a disagreement about the facts of any one payroll ledger or any one national wage series. It is a reminder that “skill” as a variable means something different — a wage-premium category, a specific task reassignment, or a cross-country conditioning factor — depending on which approach is producing the number, and comparing numbers across approaches without tracking that difference is a common source of overstated claims in popular accounts of automation and industrialization alike.
What this comparison is not
None of the three approaches here is a stand-in for, or an improvement on, the other two. Growth accounting is not “more rigorous” than business history because it uses more formal statistics — it answers a different question, at a different resolution, from different evidence, and loses the mechanism business history preserves. Business history is not “more concrete” than comparative economics in some way that makes it more trustworthy — a single firm’s ledger, however detailed, cannot by itself tell you whether that firm was typical. Comparative cross-country panels are not “more generalizable” in a way that makes the other two approaches provincial — a panel built from uneven historical statistics inherits every measurement problem present in its weakest national series, multiplied across every country included. The honest description of the field is that these three traditions grew up asking different questions of the same historical episodes, and the strongest published work on any single GPT episode — steam, electrification, or the general-purpose information technologies more likely to interest a reader in 2026 — routinely draws on results from all three rather than treating them as competitors.