Ask three political scientists how “strong” a state is and you will get three different research operations, not three different opinions about the same measurement. One will open a spreadsheet of country-year scores built from perception surveys. One will pull a box of tax rolls from an archive and reconstruct, decade by decade, how one particular bureaucracy actually came to collect revenue. One will draw a game tree on a whiteboard and derive, from a ruler’s and a taxpayer’s incentives, what capacity should exist in equilibrium. All three call themselves research on state capacity. None of them is a rough draft of the others.
This matters because the three traditions get compared casually and badly. A policy brief cites a Worldwide Governance Indicators percentile as though it were a temperature reading. A historian dismisses cross-country indices as “just perception data” without engaging what perception data is actually for. A formal theorist’s equilibrium gets treated as a prediction about a specific country next year, which it was never built to be. The way to compare the three honestly is not to rank them, but to ask what each one is actually built to answer, on what evidence, with what error bars, and where its explanatory reach runs out. That comparison is the subject of this article: quantitative cross-country governance indices, historical and archival case-study institutional analysis, and formal political-economy modeling, set against one another on measurement validity, causal inference, and generalizability.
What “state capacity” is even claiming to measure
Before comparing methods it is worth being explicit that “state capacity” and “governance quality” are not self-evidently the same variable, and different traditions quietly answer different questions under the same label. Francis Fukuyama’s widely cited essay on the concept argued that much of the empirical confusion in governance research traces back to unresolved disagreement about what “good government” even means — procedural conformity to bureaucratic norms, the capacity and professionalism of the state apparatus, the outputs it delivers, or the autonomy it has from political interference [6]. He proposed treating capacity and autonomy as two separate axes rather than collapsing them into one score, which is itself a judgment call about concepts, not just measurement — a reminder that before any index, archive, or model is chosen, a prior definitional decision has already been made about what is being studied.
Timothy Besley and Torsten Persson’s political-economy work offers a second, more restrictive definition: state capacity as the legal capacity to enforce contracts and property rights and the fiscal capacity to raise revenue, treated as investments a polity makes or fails to make depending on the incentives its political institutions create [4]. Peter Evans, writing from historical sociology, defines capacity relationally instead — as a state’s “embedded autonomy,” a combination of internal bureaucratic coherence and structured ties to the private sector it governs [5]. Three traditions, three working definitions. Comparing their empirical methods without noting this is comparing answers to different questions as though they were competing answers to one question.
Approach one: quantitative cross-country governance indices
The most visible and most cited output of state-capacity research is the composite governance index — a single number per country per year that claims to summarize something like “rule of law” or “government effectiveness” across the entire world.
The Worldwide Governance Indicators, produced by Daniel Kaufmann, Aart Kraay, and Massimo Mastruzzi at the World Bank, are the paradigm case. The WGI aggregate several hundred underlying variables drawn from more than thirty separate data sources — commercial risk-rating services, NGO assessments, cross-national surveys of firms and households, and public-sector expert assessments — into six composite indicators covering voice and accountability, political stability, government effectiveness, regulatory quality, rule of law, and control of corruption, for more than two hundred countries going back to 1996 [1]. The aggregation method is an unobserved components model: each source is treated as a noisy, partial signal of a latent governance construct, and the composite is a precision-weighted average of those signals, with an explicit margin of error reported alongside every country score.
V-Dem (Varieties of Democracy) takes a related but methodologically distinct path. Rather than aggregating existing indices and expert surveys built by other organizations, V-Dem fields its own network of over three thousand country experts, who independently code the same underlying questions — the openness of media censorship, the intimidation of election officials, and hundreds of others — on ordinal scales for every country-year in its coverage, in some cases back to 1789. A custom Bayesian item-response measurement model then estimates each country’s latent position while explicitly modeling and reporting inter-coder reliability and disagreement, rather than simply averaging raw scores [2].
What this approach is genuinely good at. Comparability at scale is the whole point, and it delivers it: a researcher can regress economic growth in 190 countries against a “rule of law” score and get an answer using the same instrument everywhere, updated annually. The explicit uncertainty bands (WGI) and disagreement metrics (V-Dem) are a real methodological advance over older indices that reported a single point estimate with no error bar at all — a country’s WGI percentile rank is frequently indistinguishable from a rank ten or twenty places away once the margin of error is taken seriously, a fact the methodology paper itself stresses [1].
Where its validity is genuinely contested. Bo Rothstein and Jan Teorell’s influential critique of the field is not aimed at any one index’s math; it argues the underlying concept most indices try to measure was never adequately specified in the first place. They propose defining quality of government specifically as the impartiality of institutions in exercising public authority, and point out that competing concepts — democracy, rule of law, effectiveness — get blended together in most composite indices without a clear theoretical rationale for the blend [8]. This is a measurement-validity problem in the technical sense: even a perfectly reliable instrument (one that produces the same score from the same inputs) can still be invalid (measuring something other than, or in addition to, the construct it claims to measure) if the underlying concept is contested or ill-defined. A second, more mechanical limitation is that most composite indices are built substantially from perceptions — expert perceptions, investor perceptions, household perceptions — which correlate with real conditions but also with the media environment, recent scandals, and the assessor’s own priors; the WGI methodology paper is explicit that this is a feature the indicators inherit from their source data, not a design flaw that can be aggregated away [1].
Causal inference. This is the sharpest limit. A cross-country regression using an annual governance index can show that “rule of law” correlates with GDP per capita, but the index is a summary snapshot, not a sequence of events; it cannot, by itself, establish which caused which, or rule out a shared prior cause producing both. Index-based studies routinely borrow instruments and natural experiments from elsewhere in the literature to address this, but the index itself is not a causal-identification device — it is a measurement device for a variable that then has to be plugged into some other empirical design to make causal claims survive.
Approach two: historical and archival case-study institutional analysis
The second tradition abandons breadth for depth. Instead of scoring every country against one instrument, it reconstructs, from primary administrative records, how one state’s capacity was actually built, exercised, or lost.
Peter Evans’s “embedded autonomy” study is a clean example of the genre in its comparative-case form: rather than scoring bureaucratic quality on a universal scale, he traces, through interviews and internal documentation, how state agencies in Brazil, India, and South Korea interacted with local firms and multinational capital in building computer industries during the 1970s and 1980s, arguing that success depended on a specific structural combination — internal bureaucratic coherence plus dense, reciprocal ties to industry — that a cross-country index has no mechanism to detect [5]. James C. Scott’s Seeing Like a State works at a still-longer historical range, using cases from Prussian scientific forestry to Soviet collectivization to Tanzanian villagization to show how modern states built the very capacity to see and tax their populations — standardized surnames, cadastral surveys, uniform weights and measures — and how that same legibility-building capacity produced some of the twentieth century’s worst administrative disasters when imposed without local knowledge [9]. Daron Acemoglu and James Robinson’s comparative-historical account in Why Nations Fail sits between the case study and a broader theory, using paired historical episodes — the two halves of a divided city, a colonial frontier, a critical juncture like the Black Death — to argue that “inclusive” versus “extractive” political institutions, not resource endowments or culture, set the trajectory of state capacity and economic development over centuries [7].
What this approach is genuinely good at. Mechanism. A cross-country index can tell you that country A scores higher on government effectiveness than country B; it cannot tell you how a particular ministry actually built the routine of collecting a tax, training an inspector, or enforcing a contract, because that sequence of decisions is exactly what gets averaged away when hundreds of variables are compressed into one annual number. Archival work can also detect capacity that indices systematically miss or misprice — a state that scores poorly on investor-perception indices because its formal rule of law is thin, but that in practice runs an effective informal taxation and dispute-resolution system the survey instrument was never designed to see, or the reverse: a state that scores well because its statute book is admired abroad while the administrative reality on the ground diverges sharply from the law on paper.
Where its generalizability is genuinely limited. A single case, or even a matched pair or triad of cases, cannot establish a base rate. Evans’s three-country study is compelling on the mechanism of embedded autonomy in those three states in that period; it is not, and does not claim to be, a statement about how common that mechanism is across all developing states, or what fraction of attempted state-led industrialization succeeds by that route versus fails. Case selection is also a live methodological problem in a way it is not for a full-coverage index: cases are usually chosen because the outcome is already interesting, which risks selecting on the dependent variable and inflating the apparent strength of whatever mechanism the researcher went in looking for.
Where its causal claims are genuinely strong, on their own terms. Process tracing — reconstructing the actual sequence of decisions, memos, and events inside a single case — can establish that a plausible mechanism connects cause to effect in a way a cross-country regression cannot, because it observes the intermediate steps directly rather than inferring them from a correlation between two endpoints. What it trades away is the ability to say how often that mechanism generalizes beyond the case examined.
Approach three: formal political-economy modeling
The third tradition does not measure existing states at all in the first instance. It starts from assumptions about actors’ incentives and derives, mathematically, what level of state capacity should exist in equilibrium.
Besley and Persson’s models are the reference point here. In “The Origins of State Capacity,” they build a formal framework in which a government chooses how much to invest in legal capacity (institutions that protect property rights and support market transactions) and fiscal capacity (the administrative machinery to collect taxes), where the two investments are complements: legal capacity expands the taxable private economy, and fiscal capacity funds the legal system, so investing in one raises the return to investing in the other [4]. The model’s testable implication is not a prediction about any single country but a comparative- statics claim: capacity investment should be higher where political institutions are more inclusive, where political power and economic interests are more aligned, and where public goods are valued highly enough to outweigh the redistributive risk that a stronger state could later be turned against the group that built it. Their book-length treatment, Pillars of Prosperity, extends the same apparatus to reinterpret Adam Smith’s “peace, easy taxes, and a tolerable administration of justice” as three complementary equilibrium outcomes of one underlying political settlement, rather than three independent policy choices a government can pick off a menu [3].
What this approach is genuinely good at. Explicitness and internal consistency. A formal model forces every assumption into the open — what actors want, what they know, what they can commit to — and then derives a strictly logical conclusion from those assumptions. This makes it possible to ask precisely which assumption is doing the work in a claimed result, something that is often impossible to isolate from a single historical narrative or a correlation in cross-country data. It is also the only one of the three approaches that produces genuinely falsifiable comparative-statics predictions (“capacity investment rises with political inclusiveness, holding X and Y fixed”) that a subsequent empirical study — using either an index or a case — can go out and test directly.
Where it is honest about its own limits. A model’s conclusion is only as good as its premises, and the premises are chosen by the theorist for tractability as much as for realism. Besley and Persson’s own state-capacity models assume, for instance, a specific and simplified structure of political competition and a small number of organized groups bargaining over policy; real polities have far messier coalition structures, informal power brokers, and path-dependent bureaucratic routines that do not reduce cleanly to the model’s actors. A formal model also cannot, by itself, tell you whether its assumptions describe any actual country at a given moment — that is an empirical question the model hands off to the other two traditions.
Setting the three side by side
None of the three traditions is a lesser or earlier draft of the other two; they trade the same three currencies — measurement validity, causal inference, and generalizability — in opposite directions.
Cross-country indices maximize generalizability: full coverage, updated annually, directly comparable across the widest possible set of cases. They pay for that with weaker causal traction (a snapshot correlation, not a mechanism) and a contested measurement-validity foundation, since what a composite score is actually measuring depends on a definitional choice about governance that the field has not settled [8] [6].
Archival and case-study analysis maximizes mechanism and measurement validity for the case at hand — a researcher who has read the actual ledgers and memos knows precisely what was measured and how, because they measured it themselves. It pays for that with weak generalizability, since a finding about embedded autonomy in Korean industrial policy or Prussian forestry administration is a claim about that case, extended to others only by analogy and further comparative work, not by statistical extrapolation [5] [9].
Formal modeling maximizes internal logical validity — a derived result follows necessarily from stated assumptions, and every assumption is visible for inspection. It pays for that by making no direct empirical claim about any actual country at all until its predictions are taken to data by one of the other two approaches [4].
The strongest recent state-capacity scholarship increasingly works by triangulation rather than substitution: a formal model generates a comparative-statics prediction, a cross-country index provides the broad-coverage test of whether the predicted pattern holds on average, and an archival case study checks whether the mechanism the model assumes is actually the mechanism operating on the ground in specific instances where the index result is surprising or contested. Besley and Persson’s own empirical chapters do exactly this, pairing their formal predictions with panel-data tests using governance indicators broadly similar in construction to the WGI [3]. Acemoglu and Robinson likewise move between broad historical pattern and close single-case narrative within the same book, using the paired cases not as decoration for a prior quantitative finding but as the primary evidence for a mechanism that a cross-country regression could not by itself identify [7].
What this means for reading a claim about state capacity
A few concrete distinctions can be applied whenever a claim about governance or state capacity is encountered outside a specialist journal.
A claim of the form “country X ranks in the Yth percentile on rule of law” is a fact about that particular index’s output, not a fact about the country’s institutions in some index-independent sense — a different composite built from different source surveys and a different aggregation rule can and does produce a different rank for the same country in the same year, precisely because the underlying concept is contested [1] [8].
A claim of the form “this ministry succeeded because of embedded autonomy” or “this land reform failed because it destroyed legible local knowledge” is an analytical interpretation of a specific historical case, built on real archival or field evidence, but bounded by that case unless further comparative work shows the same mechanism recurring elsewhere [5] [9].
A claim of the form “state capacity investment should rise with political inclusiveness” is a theoretical prediction derived from a formal model’s assumptions — worth taking seriously exactly to the degree those assumptions are judged to describe the situation at hand, and testable against either an index or a case [4].
None of these three is a weaker version of “real” measurement waiting to be replaced by a better index, a bigger archive, or a more elaborate model. They are different instruments built for different questions, and comparative political science has not converged on one because the questions themselves — how much capacity exists everywhere right now, how a particular capacity was actually built, and what capacity should exist given a stated set of incentives — are not the same question asked three ways. They are three questions.