Every population forecast that appears in a news article, a planning department’s slide deck, or a housing-affordability debate is downstream of a specific, documented, and auditable procedure. The number itself — “the country’s population will peak at X in year Y,” “this metro area needs Z new housing units by 2035” — is usually reported without the method that produced it. That is a loss, because the method is where the real information lives: what was assumed, what was measured, what was extrapolated, and where the forecast can be wrong. This guide walks through three linked procedures that practising demographers actually run — cohort-component population projection, stratified multi-stage sampling for migration surveys, and the demographic pipeline that turns a population forecast into a housing-unit demand number — in enough operational detail that a reader could audit a published projection rather than simply cite it.
Fact, claim, analysis, scenario, prediction
Before the mechanics: this piece distinguishes five kinds of statements throughout, and the distinction matters more here than in most technical writing, because demographic forecasting sits exactly on the seam between measurement and speculation.
Fact — a number with a specific, checkable source: the median age recorded in a census, the count of registered deaths in a given year. Vendor or agency claim — a statement an institution makes about its own product, such as a stated margin of error on a projection series. Analysis — a reasoned interpretation connecting facts, such as noting that a country’s headship rates have been falling and inferring an effect on household formation. Scenario — an explicit branch of a forecast under stated assumptions, such as the U.S. Census Bureau’s low-, main-, high-, and zero-immigration variants [3]. Prediction — a specific expectation with a horizon, a mechanism, and a disconfirmation condition, as distinct from a scenario, which does not claim to be more likely than its alternatives. Every forward-looking statement below is labeled as one of the latter two.
The cohort-component method: how a population is actually projected
The cohort-component method is the standard technique behind essentially every official population projection in use today, including the UN Population Division’s World Population Prospects and the U.S. Census Bureau’s national projections [1] [3]. Its logic is simple enough to run on paper for a single age group, and the complexity of a real national projection is just this logic repeated across every age, both sexes, and every year of the horizon.
The base population. The method starts from a population counted by single years of age (or five-year age groups) and sex, taken from the most recent census or population estimate. The U.S. Census Bureau’s 2023 national projections series, for instance, uses the official estimate of the resident population as of July 1, 2022, as its jump-off population, and it was the first Census Bureau projection series built directly on the 2020 Census [3].
Aging the cohort forward. For each one-year (or five-year) step, every age-sex cohort is
advanced by applying an age-specific survival ratio — the probability of surviving from one age to
the next, derived from a life table built out of registered deaths and mortality schedules — and
then subtracting deaths and adding net migrants assigned to that same age and sex [1].
Symbolically, for a cohort of age
where
The three components, and where their numbers come from. Fertility inputs come from registered births classified by mother’s age where vital registration is complete, and from indirect estimation methods — reverse survival applied to census age structures, “own-children” methods applied to census and survey microdata, and cohort-completed fertility backdated by mean age at childbearing — where it is not [2]. Mortality inputs come from registered deaths where a country has a functioning vital registration system, and from indirect demographic estimation techniques (sibling survival methods, census-based child mortality estimation) where it does not [2]. Migration is the weakest-measured of the three components in almost every national statistical system, which is the reason the next section exists: where births and deaths are typically registered as legal events, cross-border and internal moves usually are not, and have to be estimated from surveys, residual methods, or administrative proxies such as visa and residence-permit records [1].
Why scenarios, not a single number. Because the mortality, fertility, and especially the migration components cannot be forecast with certainty, official projections are published as explicit scenario bundles rather than single point forecasts. The U.S. Census Bureau’s 2023 series publishes a main projection alongside low-, high-, and zero-immigration alternative scenarios that hold fertility and mortality assumptions fixed and vary only the migration assumption [3]. This is a scenario structure, not a claim that any one branch is the likely outcome — treating a middle scenario as “the forecast” is a common misreading of exactly this kind of output.
Migration surveys: how the flow numbers are actually collected
If mortality and fertility rest on registration systems, migration mostly rests on surveys, and the survey’s sampling design determines what the resulting numbers can and cannot support. This section walks through how a household migration survey is actually built, using the design standard behind programs such as the Demographic and Health Surveys and the sampling guidance UNHCR and ODI publish for field use [7] [8].
Stage one: stratification. The target country or region is first divided into strata — typically crossing administrative region with an urban/rural indicator — because migration propensity, access, and survey cost differ systematically along exactly those lines. Sampling separately within each stratum, rather than treating the whole population as one pool, ensures that small but distinct subpopulations (a remote rural district, a border region with unusually high out-migration) are not swamped by sample drawn disproportionately from easier-to-reach urban strata [7].
Stage two: selecting primary sampling units. Within each stratum, the survey selects primary sampling units (PSUs) — census enumeration areas, villages, or urban blocks — using probability-proportional-to-size sampling, so that a PSU with more households has a proportionally higher chance of selection. This is the standard first stage of the stratified two-stage cluster design used across Demographic and Health Surveys implementations [7].
Stage three: listing and selecting households. Once a PSU (an enumeration area) is selected, field teams conduct a fresh household listing — walking the area and recording every household, because the last census’s household list is by definition already out of date — and then draw a fixed number of households, typically on the order of 25 to 30, from that fresh list by equal probability sampling [7]. This is the point at which the enumerator’s tablet or paper roster becomes the unit of data collection: each selected household is visited, its members enumerated, and migration-history questions (has anyone in this household moved in the past 12 months, does the household include anyone who migrated internationally, does the household receive remittances) are asked of a designated respondent.
Why the design matters for interpreting the output. Because selection probabilities differ by stratum and by PSU size, every respondent’s answer must be weighted by the inverse of its selection probability before national or subnational estimates can be built — an unweighted tabulation of a stratified cluster sample systematically misrepresents the population. This is also why cluster samples require a design effect adjustment to standard errors: households within the same PSU tend to resemble each other (same local labor market, same local shock exposure), which reduces the effective sample size relative to a simple random sample of the same nominal size. A migration estimate published without a stated standard error or without disclosing the sample’s clustering is missing the information needed to judge its precision.
From population to households to housing units
Neither a UN-style national population projection nor a migration survey answers the question a city planning department actually needs answered: how many housing units will this metro area need in ten years? That requires two further conversions, both of which are demographic in nature.
Population to households: the headship-rate method. The classical technique, documented in UN methodological manuals since the 1970s, applies an age-, sex-, and sometimes race/ethnicity-specific headship rate — the proportion of people in a given demographic cell who head their own household — to each cell of a population projection, then sums the resulting household counts [4]. If a projection shows 4 million people aged 30–34 and the headship rate for that cell is 0.42, the method attributes 1.68 million households to that cell alone. The method’s acknowledged weakness is that headship rates are themselves social outcomes, sensitive to housing cost, marriage timing, and the propensity of adult children or unrelated adults to double up in a single dwelling, so a projection has to make an explicit assumption about how headship rates will trend, not only about population size.
The extended cohort-component alternative. Because the headship-rate method treats “household formation” as a static rate applied to a population count rather than as its own demographic process, researchers including Zeng, Wang, and Gu have developed an extended cohort-component approach — the ProFamy family of models — that projects household and living-arrangement transitions (marriage, divorce, widowhood, children leaving home) directly, alongside the population itself, at the subnational level [5]. This produces a household forecast that is internally consistent with the demographic transitions producing it, rather than a population forecast with a headship rate multiplied on afterward — a meaningful improvement where headship rates are known to be shifting quickly, at the cost of requiring much richer input data on household transitions than most subnational statistical offices routinely collect.
Households to housing units: reconciling against existing stock. The final step is not purely demographic. The Joint Center for Housing Studies’ national household growth projections combine the demographic household forecast with assumptions about the rate of removal of existing housing units (demolition, conversion, disaster loss) and the target vacancy rate for a healthy market, to arrive at a net new construction requirement [6]. State and local agencies run a parallel version of this same pipeline: Massachusetts’s population-and-housing projection methodology, for example, documents exactly this same population-to-household-to-unit chain used to inform its own regional housing targets [11]. In practice, this is where a planning department’s forecast is reconciled against a housing-permit ledger — the actual count of units approved and built since the last projection — because permits are observed administrative facts, while the projected household count is a model output, and the two are compared explicitly rather than treated as automatically consistent.
Aging populations and the demographic transition
Population aging is a direct, mechanical consequence of the same cohort-component arithmetic run over several decades under falling fertility and rising survival: fewer young cohorts enter at the bottom of the age pyramid while more people survive into older age at the top. The UN’s 2024 population ageing policy brief documents that this is now a broad-based phenomenon rather than one confined to high-income countries — ten of today’s least-developed countries already had more than a million people aged 65 and over in 2023, and the UN projects that number will rise to 27 least-developed countries by 2050, including two with more than 10 million older people [9]. Global life expectancy at birth reached roughly 73.3 years in 2024, and further mortality decline is projected to push it toward roughly 77.4 years, a fact distinct from any single country’s trajectory and itself a summary statistic sensitive to how it is weighted across countries [9].
This matters for the housing pipeline above for a specific, non-obvious reason: aging shifts headship rates and household size independently of total population size, because older households are more likely to be single-person or couple-only households even where the total population is flat or declining. A city projecting flat population growth cannot therefore assume flat housing demand; if its age structure is aging in place, headship-rate-weighted household counts can keep rising for years after the raw population count peaks — an analytical point, not a fact about any specific place, and one that depends on that place’s actual headship-rate trajectory being verified locally rather than assumed.
Climate-linked mobility: where the method is least mature
The IPCC’s Sixth Assessment Report states, with high confidence, that human-induced climate change has already affected human mobility patterns, through both changes in migration destination choices and increased displacement risk [10]. But climate-linked mobility estimates are methodologically the least mature component discussed in this guide, and the gap is worth stating plainly rather than glossing over. Existing hazard-exposure estimates combine ensembles of climate-impact models projecting hazards — heat waves, drought, wildfire, river flooding, tropical cyclones, and crop failure — under alternative emissions pathways, and overlay those hazard projections onto population-density layers to estimate exposure [10]. That overlay step produces an exposure estimate, not a migration estimate: exposure to a hazard is not the same as displacement or migration in response to it, because whether an exposed population actually moves depends on adaptive capacity, local infrastructure investment, and policy choices that the hazard models do not represent. The IOM’s own account of this literature is explicit that most existing data on migration linked to slow-onset hazards such as drought are qualitative and case-study-based, with few comparative, quantitative studies bridging hazard exposure to observed population movement [10]. Widely cited totals — such as estimates that over a billion people could be exposed to coastal climate hazards by 2050 — are exposure scenarios under stated assumptions, not migration predictions, and treating them as forecasts of how many people will actually relocate substitutes an easier-to-model quantity (exposure) for the harder one the public debate actually wants (movement).
An explicit prediction, stated on the guide’s own terms
Prediction: over the next fifteen years, subnational household-formation forecasts built on the static headship-rate method will increasingly diverge from observed household counts in aging, high-housing-cost regions, because those forecasts hold rate trends fixed or trend them linearly against a headship-rate reality that responds nonlinearly to cost and to the age-structure shifts described above. Horizon: by roughly 2040. Assumptions: fertility continues its currently observed regional trajectories without a large discontinuous reversal, and no major policy intervention (large-scale subsidized construction, direct household-formation subsidies) changes headship-rate dynamics independently of the demographic drivers modeled. Observable indicator: a widening, published gap between headship-rate-method projections and extended cohort-component or administratively-observed household counts in the specific regions where this comparison is publicly reported, such as the Joint Center for Housing Studies’ periodic revisions [6]. Disconfirmation condition: if headship-rate-method forecasts and extended cohort-component or observed household counts remain within their historically typical margin of error over that period, or if the gap does not systematically favor one method’s direction of error, the prediction is false.
What auditing one of these numbers actually requires
None of the three pipelines above is a black box in the sense the public discourse often treats them as being. A cohort-component projection can be checked cohort by cohort against the survival ratios and fertility rates it published as inputs [1] [3]. A migration survey’s precision can be checked against its disclosed stratification, cluster design, and design effect [7] [8]. A housing-demand number can be checked against the headship-rate assumptions and vacancy-rate targets a housing agency actually used, and reconciled against the observed permit ledger [4] [6] [11]. Where the pipeline currently breaks down — climate-linked mobility chief among the cases surveyed here — the honest response is not to discard the estimate but to be precise about which stage of the pipeline is measured and which is still a scenario resting on an exposure model rather than an observed or even a well-validated behavioral response [10]. The discipline that separates a demographic forecast worth citing from one that is not is exactly the discipline this guide has tried to make visible: state the method, state the inputs, and say plainly which numbers are counted, which are estimated, and which are assumed.