The city that could not replace itself
For most of the period in which Europeans kept parish records, the largest towns buried more people than they baptised. Population still rose, which means the arithmetic was closed from outside. The identity is trivial but it is worth writing down, because almost every disagreement in this field is a disagreement about one of its terms:
Population change in a city equals natural increase, births minus deaths, plus net in-migration. The urban graveyard claim is the specific and testable assertion that for long stretches the first bracket was negative while the sum was positive: cities grew, and grew only because the countryside kept sending replacements.
The evidence for the pattern in England is unusually good. Davenport’s synthesis of English urban mortality dates the excess of burials over baptisms as especially pronounced across roughly 1650 to 1770, and puts numbers to how bad the environment was: London-born Quakers in the first half of the eighteenth century had life expectancies at birth as low as twenty-one years, and London infant mortality reached average rates as high as thirty-five per cent of live births [1]. Cities of that period were not merely unhealthy relative to villages. They were, on those numbers, demographic sinks.
The important part of the same work is the reversal. By the middle of the nineteenth century London’s life expectancy was only about four years below the national average, and large industrial cities that had every reason to be worse — Liverpool, Manchester — had become capable of natural increase, with births exceeding deaths [1]. Migration did not stop; migrants were still the majority of adults in the fast-growing manufacturing cities, and in 1851 fifty-four per cent of Londoners aged twenty and over had been born elsewhere, against seventy-seven per cent in Liverpool and seventy-two per cent in Manchester [1]. But the city no longer needed them merely to hold its own.
Against this stands a serious and long-standing objection. Allan Sharlin argued in 1978 that the observed natural decrease was substantially an artefact of who was being counted: early modern cities attracted large numbers of unmarried young migrants, whose deaths appeared in the burial register while their absent births did not appear in the baptismal one, so the resident population could have been reproducing itself while the aggregate books showed a deficit [2]. This is not a fringe position and it has not been simply refuted. Davenport’s own account leans heavily on migrant age structure and on selective return migration — the tendency of the sick, particularly the tuberculous, to go home to die — as forces that distort what the urban registers appear to say [1]. The honest summary is that the direction of the effect is well supported and its magnitude is contested, because the denominator was never observed as cleanly as the numerator.
The contrast with the present is what makes the question more than antiquarian. Jedwab and Vollrath describe historical “killer cities” that depended on in-migration, and set against them today’s low-income cities, which grow through in-migration and their own natural increase at once; their estimates suggest the postwar urban mortality transition could have doubled both the urbanization rate and the size of informal urban areas in their sample between 1950 and 2005, with roughly a third of that coming from the direct attractiveness of lower urban mortality and the remainder from population pressure pushing people into informal settlement [3]. Cities stopped killing their residents faster than they were born, and the shape of global urbanization changed as a result.
What the nineteenth century changed, and the argument about why
That reversal is the mortality transition, and its causes are genuinely disputed. The positions are worth stating as their holders state them, because the dispute is often flattened into a slogan.
Thomas McKeown’s thesis, which dominated the field for a generation, held that the decline was driven by rising living standards and above all by improved nutrition, with curative medicine contributing little. Simon Szreter’s counter-argument is that this reading omits organised social intervention: he credits the local public health apparatus — medical officers of health, sanitary inspectors, housing officers, health visitors, trained midwives, school medical officers — and insists that public health effort comes ultimately from the political realm, so that any explanation confined to income and diet mistakes the mechanism [4]. Szreter’s own position has in turn been criticised for how well it fits the infant mortality evidence, which is exactly the kind of unresolved detail that should make a reader suspicious of confident summaries.
A third strand, associated with Samuel Preston’s cross-country work and developed by Cutler, Deaton and Lleras-Muney, treats income growth as necessary but grossly insufficient to explain the size of the twentieth-century gains, which added close to thirty years of life expectancy at birth in high-income countries [5]. Their review is explicit that no single theory currently accounts for the historical decline, the cross-country differences and the within-country gradients simultaneously [5]. That is not a rhetorical hedge. It is the state of the field.
So the live question is not whether water mattered. It is how much of the decline water and sanitation can carry, once income, nutrition, housing, milk supply, vaccination and the retreat of particular pathogens are given their due. Davenport’s account of the English case is a useful corrective here, because two of the largest shifts she identifies — the collapse of smallpox mortality after vaccination spread from about 1800, and a rise in early childhood mortality from scarlet fever after 1830 — have nothing to do with the water supply at all [1].
Snow, accurately
The Broad Street pump is the most retold story in public health and one of the most misremembered. What Snow actually assembled in 1854 was two very different bodies of evidence.
The first was the Soho outbreak, a spatial argument built on a cluster of deaths around one pump — Snow wrote that within two hundred and fifty yards of where Cambridge Street joined Broad Street there were upwards of five hundred fatal attacks of cholera in ten days [8] — and on the exceptions that made the cluster informative: a workhouse and a brewery inside the affected area with their own water and conspicuously few deaths, and a widow in Hampstead, far outside it, who had pump water carried to her and died. The second, and by Snow’s own weighting the stronger, was the South London comparison, in which two companies supplied intermingled houses in the same streets — the Southwark and Vauxhall Company drawing sewage-contaminated Thames water, the Lambeth Company having moved its intake upstream. Snow and his assistant documented the circumstances of the deaths of 334 people in the first four weeks of the epidemic, going house to house [9].
What that evidence established at the time is narrower than the legend. Snow published before he had the one number the design required: the count of houses supplied by each company in the districts of mixed supply. His contemporary Edmund Parkes noticed on a second reading that Snow had therefore not actually exploited the natural experiment he had identified, and that comparing all customers of the two companies was open to confounding, because the districts differed in elevation, income and housing quality [7, 6]. Parkes’ verdict was that Snow had made waterborne transmission a hypothesis worthy of inquiry, not that he had proved it [7]. Eyler’s assessment is that Snow’s colleagues were more impressed by his evidence than by his conclusions, and that their reservation was less about the water than about his exclusiveness — his insistence that cholera entered only by water and not at all by air [6].
The pump handle did not end the argument, and Snow did not claim it ended the outbreak. Writing the following year, he recorded that the handle was removed the day after he spoke to the parish board, and then conceded that the attacks had so far diminished before the use of the water was stopped that it was impossible to decide whether the well still held the cholera poison in an active state [8]. The single most repeated intervention in the history of public health was, by its author’s own account, of undetermined effect. The General Board of Health’s Committee for Scientific Inquiries conceded that sewage-contaminated water contributed to the epidemic and still rejected Snow’s mechanism, concluding instead that the exciting cause of cholera brewed its poison from air or water containing organic impurity [6]. This is the crucial and usually omitted point about how consensus actually moved: the miasma programme did not collapse when confronted with waterborne evidence, it absorbed water as one more vehicle for morbid matter, and thereby protected itself. Strictly, what South London had established was that customers of a company supplying heavily contaminated water were at greater risk of dying of cholera. Nobody had shown how that water acted [6].
The shift came later and through other hands. William Farr, initially unpersuaded and himself the author of the elevation-based analysis that his contemporaries found more convincing than Snow’s, became the waterborne theory’s leading advocate after the 1866 East London epidemic, when his analysis pointed early and unambiguously at the East London Waterworks Company’s field and the subsequent inquiries found the company had been illegally drawing water from a contaminated reservoir at Old Ford [6]. Only with the bacteriological work of the following two decades did the mechanism acquire an agent. The lesson usually drawn from Broad Street — that a decisive natural experiment converts a scientific community — is close to the opposite of what happened.
The interventions the numbers actually support
The strongest quantitative evidence concerns two specific twentieth-century technologies: rapid sand filtration and chlorination of municipal supplies. Both were adopted city by city at different dates, which is what makes them tractable.
Cutler and Miller exploited that staggered adoption across major American cities and attributed to clean water technologies nearly half of the total mortality reduction in those cities in the first four decades of the twentieth century, about three-quarters of the infant mortality reduction and nearly two-thirds of the child mortality reduction, with a social rate of return above twenty-three to one and a cost of roughly five hundred United States dollars per life-year saved in 2003 prices [10]. Those are the numbers most often quoted, and they should be quoted with what happened to them next.
Anderson, Charles and Rees reexamined the same class of evidence for twenty-five American cities from 1900 to 1940, adding sewage treatment and bacteriological milk standards. They found water filtration associated with an eleven to twelve per cent reduction in infant mortality, and concluded that none of the other interventions they studied — chlorination included — appeared to have contributed to the observed declines; the estimated chlorination effect, on their account, was sensitive to specification and sample, losing significance under region-by-year fixed effects, without population weighting, or in levels rather than semi-logs [11]. Cutler and Miller replied in the same issue, accepting that data errors had to be corrected and reporting that their revised estimates still had filtration explaining thirty-eight per cent of the total mortality decline in their sample cities and years, against forty-three per cent originally, while conceding that the infant mortality effects were smaller than they had first reported; they attribute much of the remaining gap to the coding of partial intervention years and to the choice of population denominators [12].
This is what a live dispute looks like when both sides are competent. The disagreement is not about whether clean water reduces mortality. It is about magnitude, about which of the two technologies carries the effect, and about whether the aggregate mortality series can bear estimates that precise. My reading of the exchange is that filtration survives scrutiny better than chlorination does, and that anyone quoting a single headline percentage without the correction and the counter-estimate is misrepresenting the literature.
Independent designs point the same way with different numbers. Ferrie and Troesken, studying Chicago from 1850 to 1925, where the crude death rate fell by sixty per cent, attribute between thirty and fifty per cent of that decline to the introduction of pure water, and find much smaller effects for diphtheria antitoxin and milk inspection [14]. Alsan and Goldin, using the staged construction of the Boston-area water and sewerage districts in Massachusetts from 1880 to 1920, find that clean water and effective sewerage were complementary rather than substitutable and together account for about one-third of the decline in log child mortality over those forty-one years — with the explicit implication that piecemeal infrastructure is unlikely to move child health much [13]. Kesztenbaum and Rosenthal reach a compatible conclusion for Paris between 1880 and 1914, where the diffusion of the sewer network across the city’s eighty neighbourhoods contributed materially to life expectancy, on the argument that the benefit of clean water depends on having a means of removing the waste water afterwards [15].
The complementarity result matters more than any individual coefficient, because it explains why single-intervention estimates disagree. If supply and disposal only work together, then a study that observes one arriving without the other will measure a small effect, and both studies can be right about their own city.
Somebody has to be able to borrow
Underneath the epidemiology sits a problem that is fiscal before it is medical, and it is the part of the story that most accounts skip.
A waterworks is long-lived capital. It has to be paid for in a lump, before anybody’s child survives, and repaid out of rates collected for decades afterwards from people who cannot easily be excluded from the benefit. That combination — high fixed cost, long horizon, non-excludable returns — is not a technical problem, it is a financing problem, and it requires an institution capable of committing future revenue.
Cutler and Miller’s account of American municipal waterworks makes the point historically. The first large-scale municipal system was completed in 1801, yet many American cities had no waterworks until the turn of the twentieth century, and construction then arrived in a rush from 1890 through the 1920s. Their explanation privileges the development of local public finance, and specifically the growth of the municipal bond market as an enabler of debt finance, over alternatives including knowledge of disease transmission, externalities, density, natural monopoly, contracting difficulty, corruption and the availability of engineers [16]. This is one authors’ argument rather than a settled finding, and it should be read as such. But it reframes the question productively: for a century the germ theory was not the binding constraint, and neither was the engineering. The binding constraint was the ability to borrow.
The reframing generalises. It suggests that the historical sequence usually presented as discovery, then persuasion, then construction was closer to construction becoming financeable, then being justified by whatever theory was current. It also predicts where the problem should persist today: in jurisdictions whose utilities cannot raise long-term capital because tariffs do not cover costs, collection is weak, or the borrowing entity lacks the legal standing to pledge future revenue. That is a testable claim about the present, not only a reading of the past.
Why infant mortality is the sensitive instrument
Nearly every credible study in this literature reaches for infant and child mortality rather than the crude death rate, and the reason is mechanical rather than sentimental.
Infants have a high ratio of water turnover to body mass, so a fluid loss that an adult absorbs is fatal to them. They have no acquired immunity to enteric pathogens their parents met years earlier. They consume water they did not fetch and cannot inspect, mediated through weaning foods and feeding vessels, which means household water quality reaches them undiluted by choice. And they are concentrated in a narrow age band, so a change in the water supply shows up as a change in a rate rather than being smeared across a lifetime of accumulated exposures. Deaths at older ages reflect decades of nutrition, occupation and previous infection; deaths in the first year reflect mostly the last few months of environment.
That sensitivity is exactly why the disputed estimates cluster there. Cutler and Miller’s largest original claim was about infant mortality, Anderson, Charles and Rees’ surviving positive result is about infant mortality, and Cutler and Miller’s concession in their reply was specifically that the infant mortality effects shrank [10, 11, 12]. An instrument sensitive enough to detect the signal is also sensitive enough to move when the specification does. Ferrie and Troesken’s Chicago decline is likewise driven by infectious disease and by infant and child mortality [14].
The problem did not end, it moved
Treating this as settled history is the most common error. The WHO/UNICEF Joint Monitoring Programme reports that in 2024, 2.1 billion people still lacked safely managed drinking water services — of whom 1.4 billion had a basic service, 287 million a limited service, 302 million an unimproved source, and 106 million were drinking surface water — while 3.4 billion lacked safely managed sanitation and 1.7 billion lacked even basic hygiene services [17].
The definitions are the argument. Under the SDG ladder, a source counts as safely managed only if it is improved, accessible on premises, available when needed, and free from faecal and priority chemical contamination; a source is merely basic if the round trip to collect water, including queuing, takes no more than thirty minutes, and limited if it takes longer [17]. Availability and queuing are inside the definition, which is the modern restatement of the nineteenth-century standpipe: a tap that runs four hours a day supports a different household than one that runs continuously, even where the water leaving both is identical.
The urban numbers are the ones that should trouble anyone who assumes growth solves this. Between 2000 and 2024, 1.5 billion urban residents gained safely managed drinking water, and urban coverage stayed flat at eighty-three per cent, because the urban population grew by two-thirds over the same period; in urban areas the number of people spending more than thirty minutes on a round trip to collect water from an improved source doubled, from 44 million to 90 million [17]. Enormous absolute progress, no proportional gain, and a growing minority pushed back down the ladder. WHO’s drinking-water fact sheet, in the version last updated in September 2023, puts the associated burden at about one million deaths a year from diarrhoea attributable to unsafe drinking water, sanitation and hand hygiene, including 395,000 deaths of children under five that could be avoided if those risk factors were addressed [18].
What would change my reading
Two things follow that are worth stating as claims with failure conditions rather than as conclusions.
The first is analytical. If the complementarity result generalises, then programmes that deliver water without disposal should show substantially smaller mortality effects than programmes delivering both, in contemporary as well as historical data. Alsan and Goldin’s Massachusetts estimate and Kesztenbaum and Rosenthal’s Paris estimate both point that way [13, 15]. It would be disconfirmed by well-identified contemporary evidence showing full mortality benefits from water quality improvements in settings with no corresponding change in waste removal.
The second is a forecast, on a horizon of roughly a decade to 2035, and it is conditional. Assuming urban population growth in low-income countries continues near its current pace and utility financing conditions do not materially improve, I expect global urban coverage of safely managed drinking water to remain roughly flat in percentage terms while the absolute number served rises substantially, and I expect the population spending more than thirty minutes collecting water in urban areas to keep rising. The observable indicators are the JMP’s urban coverage percentage and its urban collection-time series, both published on the same basis today [17]. The disconfirmation condition is straightforward: a sustained rise in urban safely-managed coverage of more than a few percentage points over a decade, accompanied by a fall in the urban collection-time population, would show that supply is outrunning urban growth and that the fiscal constraint has loosened.
The general claim I am willing to defend is narrower than the title suggests and, I think, better supported for being narrow. Water and sanitation did not cause the mortality transition on their own; the attribution remains contested and the best estimates disagree by a factor of two or more. What water did was remove the specific mechanism by which density killed. Everything else a city offers — labour markets, specialisation, the accumulation of skill and capital in one place — was available in the eighteenth century too, and cities still could not hold their populations without a constant supply of newcomers. The demography came second. The plumbing came first.