Ask a comparative political scientist whether a given country has “high state capacity” and the honest first answer is another question: capacity to do what, measured how, compared against which cases, over what period? “State capacity” is not a quantity sitting in the world waiting to be read off an instrument the way a thermometer reads temperature. It is a research construct, assembled from proxies that were built for other purposes — tax bookkeeping, census enumeration, expert opinion surveys, legal gazettes — and pressed into service as evidence about something none of them were designed to measure directly. That gap between the construct and the proxy is not a footnote to the method; it is most of the method.
This is a practitioner’s walkthrough of three things comparative-governance researchers actually do: build a capacity index out of administrative-data proxies, stage a structured case comparison so that institutional variables carry the argument rather than the researcher’s intuition, and locate and read legal and administrative primary sources in archival practice. Each claim below is tied to a specific, checkable source. Where the claim is a vendor’s or an institution’s own description of its product, that is marked. Where it is my analytical inference rather than a documented finding, that is marked too.
What “state capacity” is asked to explain
Francis Fukuyama frames the underlying distinction cleanly: a state’s capacity — its ability to make and carry out decisions — is conceptually separate from its size and from whether it is democratic or accountable [8]. Singapore and the Netherlands can both be judged highly capable states while differing enormously in the scope of what they do; a large state apparatus is not automatically a capable one. Fukuyama’s own tripartite framework — a capable state, the rule of law, and democratic accountability, all three needed together and none substituting for the others — is itself an analytical claim about political development, not a measurement protocol, and I flag it here as analysis rather than as an empirically settled fact.
Timothy Besley and Torsten Persson give the concept sharper analytical edges by splitting it into two complementary components that a government invests in over time: fiscal capacity, the ability to raise revenue with a broad, low-distortion tax base, and legal capacity, the ability to support markets by enforcing contracts and property rights [2]. Their model’s finding — a fact about the model, not a description of any one country — is that these two capacities are typically complements: investment in one raises the return to investing in the other, and both are more likely to be built where there is political stability, inclusive institutions, and common external threats such as war that make the investment collectively worthwhile [2]. That is a theoretical result derived from a formal model calibrated against cross-country data, and it should be read as a structured hypothesis about why capacity varies, not as a direct measurement of capacity itself.
Building an index from administrative-data proxies
None of the widely used capacity indicators asks “how much state capacity does this country have?” as a survey question. Instead, researchers assemble proxies from records that already exist for unrelated administrative reasons, and each proxy family has its own failure modes.
Fiscal proxies. The most direct capacity proxy is the tax take: total revenue collected relative to GDP, ideally split by tax instrument (trade taxes, which are easy to collect at a handful of border points, versus broad-based income and value-added taxes, which require identifying and monitoring millions of taxpayers). The ICTD/UNU-WIDER Government Revenue Dataset exists specifically because standard cross-country revenue figures were inconsistent and hard to compare: it was built between 2010 and 2014 to produce a more complete and accurate database of government revenues for 196 countries, reporting figures both inclusive and exclusive of natural-resource revenue so that oil- and mineral-rich states are not misread as fiscally capable when they are merely resource-rich [4]. In practice this means a researcher building a fiscal-capacity proxy has to make an explicit choice about resource revenue before comparing countries at all — a Gulf state’s total revenue-to-GDP ratio and its non-resource revenue-to-GDP ratio can tell entirely different stories, and conflating them is one of the most common errors in cross-national fiscal comparison. The OECD’s Tax Administration series, now in its thirteenth edition and covering 58 jurisdictions using the International Survey on Revenue Administration, supplies the administrative-process layer beneath the revenue total: staffing levels, digitization of filing, audit coverage, and dispute-resolution timelines, aimed at tax officials and finance ministries who want to benchmark how revenue is actually collected rather than only how much is collected [7]. That framing — a benchmarking tool “intended to be used by tax officials,” in the OECD’s own description — is the publisher’s stated purpose, not an independent researcher’s verdict on its analytical validity, and I note it here as an institutional self-description.
Bureaucratic-structure proxies. Revenue totals say nothing about how an administration is organized to collect it. Peter Evans and James Rauch’s classic study surveyed the core economic agencies of 35 developing countries for 1970–1990 and constructed a “Weberianness Scale” scoring each agency on meritocratic recruitment and predictable, rewarding long-term careers — the two organizational features Max Weber identified as distinguishing a rational-legal bureaucracy from patrimonial administration [3]. Their finding was that higher Weberianness scores were associated with faster subsequent economic growth even after controlling for initial GDP per capita and human capital [3]. This is a genuinely useful proxy precisely because it does not ask about outcomes at all — it scores recruitment and career-structure rules, which are observable from civil-service statutes and personnel records, and treats growth as the thing to be explained rather than folding growth into the capacity measure itself. A proxy that scores organizational rules is far more defensible than one that scores outcomes and then claims capacity explains those same outcomes — that circularity is one of the most common design errors in this literature, and Evans and Rauch’s design is worth studying specifically because it avoids it.
Perception and expert-assessment proxies. The World Bank’s Worldwide Governance Indicators are the most widely cited governance measure in cross-country empirical work, and understanding what they actually are matters more than citing the headline score. The WGI aggregate six indicators — Voice and Accountability, Political Stability, Government Effectiveness, Regulatory Quality, Rule of Law, and Control of Corruption — across 214 economies from 1996 onward, built by statistically combining 35 separate underlying data sources: household and firm surveys, and expert assessments from multilateral organizations, NGOs such as Freedom House and the World Justice Project, and commercial providers such as the Economist Intelligence Unit [1]. Two features of the construction matter for anyone using the WGI as a capacity proxy. First, the underlying sources are themselves subjective perceptions, not administrative counts — the WGI is explicitly a survey-of-surveys, weighting and rescaling expert and respondent opinion rather than measuring, say, the number of tax auditors per capita. Second, and less widely appreciated, the WGI methodology paper itself reports an explicit margin of error for every country-year estimate, reflecting how many underlying sources cover that country and how much those sources agree [1]. A researcher who reports a WGI point score without its confidence interval is discarding information the data’s own publisher considered important enough to compute and release. The Fragile States Index takes a related but distinct approach: the Fund for Peace’s proprietary CAST platform triangulates data from roughly a hundred sub-indicators across cohesion, economic, political, and social categories, built partly from algorithmic content analysis of daily news and reporting on each of 178 countries [5]. That the underlying process runs through the vendor’s own proprietary software and search-term algorithms is the Fund for Peace’s own description of its method, not an independently audited protocol, and it is worth naming as a vendor claim precisely because it is harder to replicate than a public microdata release like the GRD.
A practical worked design. Putting these proxy families together, a defensible index-construction exercise for, say, a paired study of two mid-income states would proceed: (1) pull non-resource tax revenue as a share of GDP from the GRD for a common ten-year window, since resource revenue and the underlying tax effort answer different questions [4]; (2) code civil-service recruitment statutes for meritocratic entry and tenure protection, following the Evans-Rauch logic of scoring rules rather than results [3]; (3) pull the Government Effectiveness and Rule of Law components of the WGI, reporting the published margin of error alongside the point estimate rather than the point estimate alone [1]; and (4) treat the Fragile States Index as a triangulation check rather than a primary input, given its reliance on a non-public scoring algorithm [5]. This four-step sequence is my own synthesis of how these sources are typically combined in the literature, not a documented protocol drawn from any single cited study — it is presented as an analytical recommendation, and a reader should treat it as one design choice among several defensible ones, not as a standard the field has agreed on.
Structured, focused comparison: pairing cases so the institution does the explaining
A capacity index tells you where a country sits on a scale; it does not by itself tell you why. That question is usually pursued through structured case comparison — choosing a small number of cases that are matched on as many background conditions as possible so that the institutional variable of interest is the main thing left to vary. The logic descends from Mill’s method of difference, and its discipline in practice is almost entirely about case selection rather than about narrative skill: a comparison is only as strong as the argument that the paired cases really are twins except for the variable under study.
In practice this means writing down, before any archival work begins, the matching criteria (comparable population, comparable colonial or state-formation history, comparable resource endowment, comparable starting GDP per capita) and the one or two institutional variables the comparison is meant to isolate (a particular civil-service reform, a particular tax law, a particular administrative reorganization). Besley and Persson’s finding that fiscal and legal capacity co-move as complements is itself a warning for case selection: if a researcher picks two states that differ in fiscal capacity, they are also likely to differ in legal capacity for the same underlying reasons, which means a comparison naively attributing an outcome to “the tax reform” may be attributing it to a bundle of co-moving institutional changes that happened together [2]. A well-designed comparison states this risk explicitly and either finds a case where the two components diverged, or downgrades the causal claim from “this reform caused that outcome” to “this reform is associated with that outcome, and legal capacity moved together with it in a way the design cannot fully separate.”
Archival practice: reading the primary record itself
Indices and case comparisons both eventually rest on someone having read an actual administrative record — a tax statute, a civil-service ordinance, a colonial gazette, a parliamentary committee report. James Scott’s Seeing Like a State is the standard reference for why states produce these records at all: Scott’s central argument is that states pursue “legibility,” simplifying and standardizing diverse local arrangements — land tenure, forestry, naming practices — into forms a central administration can read, count, and tax, and that this drive toward legibility, when combined with an unchecked confidence in scientific planning that Scott calls “high modernism,” produced some of the twentieth century’s worst-documented failures, from Soviet collectivization to forced villagization in Tanzania [6]. This is Scott’s own interpretive argument, not a neutral inventory of facts, and other historians dispute how uniformly it applies outside his chosen cases — it is cited here as an influential analytical framework for why legible administrative records exist, not as a settled empirical law about every archive a researcher will encounter.
The practical archival workflow that follows from taking any such record seriously runs through several steps before the document itself is read: locating the record through the archive’s own finding aids (a card catalogue, an online index, or a gazette’s own contents pages, which are frequently incomplete for the periods a capacity researcher cares about most); establishing provenance, meaning which office produced the record, under what statutory authority, and what happened to the file between creation and its current location; and checking for later amendment, since colonial and post-independence administrative law is frequently amended by short subsequent ordinances that are easy to miss if only the original text is read. A tax ledger’s numbers, read in isolation from the ordinance that authorized the tax and any later ordinance that changed the rate or base, is close to meaningless for a capacity comparison across years — the researcher needs the legal instrument and the administrative execution record together, and reconciling the two is frequently where a study’s real evidentiary work happens, not in the index arithmetic that follows.
Separating claim types
To keep the layers distinct: it is a fact, checkable against the cited methodology papers, that the WGI combines 35 underlying sources and reports margins of error [1], that the GRD reports revenue both with and without resource income [4], and that Evans and Rauch scored 35 countries’ economic agencies on recruitment and tenure rules [3]. It is a vendor/institutional claim, not independently verified in this article, that the Fragile States Index’s proprietary software correctly separates relevant from irrelevant reporting [5], and that the OECD series’ benchmarking use case delivers the “dialogue among tax officials” it states as its purpose [7]. It is analysis that fiscal and legal capacity’s complementarity creates a case-selection risk for paired comparisons, and that a four-step index-construction sequence combining these proxies is a reasonable synthesis rather than a documented standard. And it is scenario/prediction, clearly bounded, that a comparative-governance dataset ecosystem increasingly dependent on algorithmically scored news content (as in the Fragile States Index’s CAST platform) will need explicit large-language-model-assisted content analysis audited against the older manual coding it replaced within the next five to ten years; the observable indicator would be published inter-coder or model-to-manual agreement statistics in the Fund for Peace’s own methodology documentation, and the disconfirming condition is the continued absence of any such published reconciliation exercise through the end of the decade.
Surveillance, coordination, and the legibility trade-off in practice
The capacity to tax and the capacity to monitor are historically bundled, and this is where the practitioner’s proxies run into their sharpest ethical and analytical tension. A cadastral survey that lets a treasury assess land tax accurately is, mechanically, the same administrative technology as a population registry that lets a security service track movement. Scott’s account of legibility is again the reference point here: the same simplifications that let a distant office read, sum, and compare — standardized surnames, uniform land parcels, single official languages — are what make a population governable in the coercive sense as well as the fiscal one [6]. A researcher coding “administrative capacity” from proxies like registry completeness or ID-card coverage should therefore treat a high score as ambiguous by construction: it is equally consistent with a state that delivers public goods efficiently and with one that surveils its population efficiently, and the proxy alone cannot distinguish the two readings. Separating them requires bringing in a second axis entirely — typically drawn from rule-of-law and accountability measures such as the WGI’s own Voice and Accountability and Rule of Law components — and reporting capacity and constraint as a pair of scores rather than folding them into one number [1]. Collapsing “capable” and “constrained” into a single composite index is one of the more common design errors in practitioner work, because it produces a single ranking where a two-dimensional one is what the underlying theory actually calls for.
Coordination failure is the mirror-image problem. A state can score well on every proxy above at the national level while individual line ministries fail to coordinate with each other or with subnational governments, producing exactly the kind of implementation gap that a purely top-down capacity index cannot see. The practical fix in fieldwork is to supplement national-level proxies with agency-level or program-level administrative records — procurement timelines, inter-agency memoranda, budget execution rates by ministry — which is precisely the kind of material that shows up only in the archival record and almost never in a cross-country dataset. This is one reason the archival and quantitative halves of this method are not alternatives to each other; a capacity claim built from index scores alone, without a spot-check against the administrative record for at least one or two of the cases in the comparison, is a claim resting on proxies the researcher never verified against the underlying paper trail.
Practical limits worth stating plainly
Every proxy family here measures something adjacent to capacity, not capacity itself. Tax revenue is shaped by tax policy choices as much as by collection capability; a state can have excellent collection machinery and still show low revenue-to-GDP because it has chosen low rates. Bureaucratic-structure scores like Evans and Rauch’s are only as good as the civil-service statutes on paper actually matching practice on the ground, which is exactly the gap that legal-archival verification work exists to check. And perception-based indicators like the WGI inherit whatever biases the underlying expert pool carries — a point the WGI’s own authors address by publishing margins of error rather than by claiming the aggregate is bias-free [1]. None of this is a reason to abandon these proxies; it is a reason to report them with their documented uncertainty, to state explicitly which component of “capacity” each proxy is standing in for, and to treat the archival record as the check against which the index is periodically recalibrated rather than as a redundant afterthought.