Three questions that are usually asked as one

After a destructive heatwave, flood or windstorm, the question that arrives first is whether climate change caused it. That question has no answer in the form it is asked, and the reason is not evasion. Any specific weather event is the product of a particular arrangement of the atmosphere on a particular week, so almost any event could have occurred by chance in an unmodified climate. The 2016 National Academies assessment put the point plainly: a definitive answer to whether climate change caused a particular event cannot usually be given in a deterministic sense, because natural variability almost always plays a role [7]. The IPCC’s Sixth Assessment repeats the same caution in its public-facing material, stating that scientists cannot answer directly whether a particular event was caused by climate change, and that what they can do instead is quantify the relative importance of human and natural influences on that event’s magnitude or probability [1].

Underneath the popular question sit three technical ones, and conflating them is the most common failure in reading attribution results.

The first is detection: has a statistically significant change occurred in some property of the climate record, beyond what internal variability would produce? This is a question about data, and it can be answered without naming a cause.

ADVERTISEMENT

The second is attribution of a detected trend to forcing: given that a change has been detected, how much of it is consistent with anthropogenic forcing and inconsistent with natural forcing and internal variability alone? Trend detection using optimal fingerprinting is described by AR6 as a well-established field, though one that meets specific difficulties when applied to extremes, because the method generally requires approximately Gaussian data and extremes are by construction not Gaussian [1].

The third is event attribution: how has anthropogenic forcing changed the probability or intensity of an event of this class, defined by a threshold, a region and a duration? AR6 dates the emergence of this as a distinct research programme to the period after the Fifth Assessment, and describes the commonly used version as a probability-based approach that produces statements of the form that climate change made this event type twice as likely, or made it fifteen per cent more intense [1].

These are separable in practice. A trend may be detectable without being attributable, and an individual event can be analysed in a region where the local trend is not detectable at all, because the model ensemble supplies statistical power the short observational record does not. Lloyd and Shepherd make the logical structure explicit: the question of what effect anthropogenic climate change had on a phenomenon can be posed unconditionally, treating everything else as a potential confounder, or conditionally, treating some causal factors as mediators to be held fixed. They argue that these framings look identical in ordinary language but admit different valid answers, and that aggregated statements about a population of events cannot be transferred reliably to an individual case [6].

The counterfactual is built, not remembered

Event attribution rests on a comparison between the world as it is and a world without anthropogenic forcing. The second world is not a historical period. It is a model construction: an ensemble of simulations run with pre-industrial greenhouse gas concentrations, or with the observed warming trend statistically removed from a fitted distribution. AR6 describes the procedure as estimating probability distributions of an index characterising the event in today’s climate and in a counterfactual climate, then comparing either intensities at fixed probability or probabilities at fixed magnitude [1].

That construction carries the analysis. Everything downstream — the ratio, the confidence interval, the headline number — inherits whatever the counterfactual got wrong. This is why the World Weather Attribution protocol devotes an entire step to model evaluation before any attribution number is computed, and why its authors set out five feasibility questions to be answered before a study is attempted at all, including whether sufficient historical observations exist and whether available models can represent the extreme in question [3].

ADVERTISEMENT
Two adjacent stone stilling wells below a harbour wall, one inlet open and flowing, the other being closed by a brass blanking plate that is only half seated, with water still slipping past its edge
Figure 1. A counterfactual world is built rather than remembered, as a second apparatus deliberately cut off from the forcing that drives the first.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

The protocol also formalises a step that sounds procedural and is in fact substantive: the event definition must specify a physical variable, a spatial extent and a temporal scale, and those choices should be tied to the impact where feasible [3]. Van Oldenborgh and colleagues, writing from within the same group, list event definition and the selection bias it invites first among the pitfalls of the field [5]. AR6 makes the sensitivity quantitative in direction if not in magnitude: attribution statements depend on the spatial and temporal extent of the event definition, and large-scale averages generally yield higher attributable changes because averaging smooths out noise [1]. An analyst who defines an event broadly will, other things equal, report a larger effect than one who defines it narrowly. Neither is cheating. The number is simply not a property of the weather alone.

Reading a risk ratio without over-reading it

The standard output is a probability ratio, also called a risk ratio, comparing the annual probability of the event in the factual climate with its probability in the counterfactual one [3]. Writing those probabilities as the factual and counterfactual annual exceedance probabilities,

PR=p1p0,FAR=1p0p1=11PR. \mathrm{PR} = \frac{p_1}{p_0}, \qquad \mathrm{FAR} = 1 - \frac{p_0}{p_1} = 1 - \frac{1}{\mathrm{PR}}.

The fraction of attributable risk is a monotone transform of the ratio and contains no additional information. Its name invites a misreading that the field has warned about for two decades: it is not the probability that this event was caused by forcing, and it does not license the statement that a given fraction of the damage was anthropogenic in any individual case. It is a population-level quantity, and Lloyd and Shepherd’s central methodological point is that population-level attributions do not transfer to singular events without further argument [6].

The arithmetic also misbehaves exactly where the events are most interesting. As the counterfactual probability falls toward zero,

limp00PR=,limp00FAR=1. \lim_{p_0 \to 0} \mathrm{PR} = \infty, \qquad \lim_{p_0 \to 0} \mathrm{FAR} = 1.

AR6 records that several studies of events from 2016 onward reported an infinite risk ratio, corresponding to a fraction of attributable risk of one, because the event’s occurrence probability was close to zero in simulations without anthropogenic influence — while noting in the same passage that the lower bound of the uncertainty range is difficult to estimate accurately in these cases [1]. Paciorek, Stone and Wehner develop a frequentist framework for exactly this sampling problem and argue that the bootstrap widely used in the literature performs poorly for small estimated probabilities, recommending methods with better behaviour in repeated samples [11].

Two small brass tipping buckets hung under the drains of a tide house, the left one caught at the instant of balance with its water bulging over the lip and not yet spilled, the right holding a single unshed drop
Figure 2. As the counterfactual count falls toward nothing the ratio between the two runs away without bound, which is why only its lower bound stays defensible.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

The practical consequence is that the lower bound is usually the defensible number. AR6 observes that while best estimates of risk ratios carry large uncertainty, their lower bounds can be relatively insensitive to observational and model uncertainty, which is what supports conservative attribution statements [1]. This is why careful studies phrase results as “at least” rather than as a point estimate. The Pacific Northwest heatwave analysis of June 2021 concluded that the event was at least 150 times less likely without human-induced climate change, and separately that it was about two degrees Celsius hotter, with a 95 per cent interval of 1.2 to 2.8 degrees, than it would have been in an 1850 to 1900 baseline climate [4].

ADVERTISEMENT

Magnitudes vary enormously by region for physical rather than methodological reasons. AR6 reports risk ratios on the order of one hundred for heat extremes in the Mediterranean and western Europe, and much less pronounced changes in the United States, attributing the difference partly to land-surface feedbacks that enhanced 1930s temperatures and so reduced the apparent rarity of recent extremes, and partly to differences in event definition and framing [1]. A cross-country ranking of risk ratios drawn from studies with different event definitions would therefore be meaningless, and the field does not construct one.

Two families of method, and a disagreement that is still live

The probabilistic family treats the specific weather pattern as a draw from a distribution and asks how the distribution moved. The storyline or conditional family holds the observed circulation fixed and asks how the thermodynamic state changed the event that actually occurred.

AR6 describes the difference as a matter of how much of the climate system is conditioned upon. Least-conditional approaches consider the combined effect of overall warming and circulation change, often using fully coupled models. More conditional approaches prescribe sea surface temperature and ice patterns, or prescribe the large-scale atmospheric circulation itself and use weather-forecasting models; the chapter identifies these highly conditional approaches with the storyline label [1].

Each family is better at something the other handles badly. AR6 states that storyline methods can be useful for events too rare to analyse otherwise, or where the specific atmospheric conditions were central to the impact, and that they enable very high-resolution simulation in cases where lower-resolution models do not represent the regional dynamics well [1]. Against that, the chapter is equally direct about the cost: the imposed conditions limit an overall assessment of anthropogenic influence, because the fixed aspects of the analysis may themselves have been affected by climate change, and a conditional hindcast permits only a statement about the magnitude of the storm had similar large-scale patterns occurred, precluding any statement about frequency if used in isolation [1].

The disagreement is not settled, and it is not merely technical. AR6 notes that the absence of model evaluation in early event attribution studies drew criticism of the emerging field as a whole, and presents the storyline approach as an alternative that does not depend on a model’s ability to represent circulation reliably [1]. A third line has developed alongside both. Faranda and colleagues analyse eight impactful 2021 events using atmospheric analogues drawn from ERA5 reanalysis over 1950 to 2021, comparing the 33 closest sea-level-pressure analogues in a 1950 to 1979 counterfactual window against a 1992 to 2021 factual window, and position the method as complementary to statistical event attribution rather than as a replacement, precisely because it foregrounds circulation dynamics that probabilistic approaches average over [10]. Their results are not uniformly attributable: several of the events they examine carry caveats about interannual and multidecadal variability that limit definitive claims [10].

Two brass pulleys carrying float wires on a whitewashed tide-house wall, the left one turning freely and the right one with a leather-faced brake shoe just meeting its rim while the wire still creeps beneath it
Figure 3. The two families of method differ only in how much of the mechanism they hold fixed, and whatever is clamped can no longer be part of the answer.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Where the confidence genuinely is

Confidence is not distributed evenly across event types, and the ordering has a physical basis rather than a sociological one.

Heat is the clearest case. AR6 assesses it as virtually certain that the frequency and intensity of hot extremes have increased and those of cold extremes decreased on the global scale since 1950, and as virtually certain that human-induced greenhouse gas forcing is the main driver of those changes globally, with very likely attribution on most continents [1]. The 2016 National Academies report explains the mechanism for this ordering: changes in long-term mean conditions provide a direct basis for expecting related changes in extremes, so temperature-related events have the strongest footing [7]. This is why the earliest formal event attribution worked. Stott, Stone and Allen chose a summer-mean temperature threshold for continental Europe that had been exceeded in 2003 but in no other year since the instrumental record began in 1851, and concluded that it was very likely, at a stated confidence level above 90 per cent, that human influence had at least doubled the risk of exceeding it [2].

Heavy precipitation sits one step down. The National Academies place hydrological drought and heavy precipitation after temperature in the confidence ordering, on the grounds that a moister atmosphere is a relatively direct consequence of warming, though less direct than the temperature change itself [7]. AR6 assesses human influence, in particular greenhouse gas emissions, as likely the main driver of the observed global-scale intensification of heavy precipitation over land, and anchors the expected scaling to the roughly seven per cent per degree increase in the moisture-holding capacity of the atmosphere [1].

Where the confidence is not, and why

Severe convective storms are the field’s honest failure. AR6 assigns low confidence to past trends in hail, convective winds and tornado activity, and gives the reason explicitly as the short length of high-quality data records [1]. The 2016 National Academies report is blunter still, finding little or no confidence in attribution of severe convective storms and extratropical cyclones, because these are governed by atmospheric circulation and dynamics that are less directly controlled by temperature, less robustly simulated by models, and less well understood [7]. AR6’s public-facing summary states without hedging that attribution of certain classes of extreme weather, tornadoes among them, is beyond current modelling and theoretical capabilities [1].

Three physical facts produce this. First, scale. AR6 notes that a tornado has a spatial scale as small as under 100 metres and a temporal scale as short as a few minutes, against a drought that can last years across a continent [1]. Second, the observing system. Diffenbaugh, Scherer and Trapp open their analysis by noting there is no reliable, independent, long-term record of severe thunderstorms — and particularly of tornadoes — with which to analyse variability and trends systematically [9]. What exists is a report archive shaped by population density, chaser networks and changing rating practice. Third, competing forcings. Convective severity requires both instability and vertical wind shear, and warming is expected to increase the first while reducing the second.

The workaround is to model the environment rather than the storm. Diffenbaugh and colleagues analyse the CMIP5 ensemble under RCP8.5 and find robust increases in days with severe thunderstorm environments over the eastern United States in all four seasons; winter shows the largest relative increase, exceeding 50 per cent by the end of the century, while spring shows the largest absolute increase, with robust increases exceeding 2.4 days per season over the central United States for 2070 to 2099, and summer the smallest relative increase at under 20 per cent [9]. They also resolve the shear objection on its own terms, finding that projected decreases in shear are concentrated on low-instability days and therefore do not reduce the total occurrence of severe environments [9].

The authors are careful about what this supports. They state that severe environments sometimes fail to produce severe thunderstorms, that CAPE, shear and convective inhibition are implicit indicators, and that mesoscale and synoptic-scale processes important for convection initiation are not captured by those indicators [9]. AR6 reaches a matching assessment from the other direction, giving medium confidence to increases in favourable spring environments and low confidence for summer, and describing how tornadoes or hail will change as an open question [1]. The observational picture is correspondingly strange: AR6 assesses with medium confidence that the mean annual number of United States tornadoes has stayed roughly constant while their variability has increased since the 1970s, with fewer tornado days per year and, with high confidence, more tornadoes on those days [1].

A deep soft coil of plain unmarked paper already wound off a clockwork drum and spilling in loops along a pale stone bench, with one fresh turn just lifting from the drum
Figure 4. Return periods are estimated from the length of the record, not from the sharpness of any single reading.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Two binding limits

Record length is the first, and it binds hardest where the science is most needed. AR6 observes that the events studied are geographically uneven, that studies in the developing world are generally lacking, and that this reflects a lack of observational data and reliable models rather than a lack of interest [1]. The 2026 National Academies assessment identifies the same asymmetry, noting that in under-resourced regions model limitations and the absence of consistent long-term records compromise attribution capacity [8].

Record length also fails at the top of the distribution even where records are long. The June 2021 Pacific Northwest event reached a peak far outside the range of historical observation and exceeded the statistical upper bound of extreme-value models fitted to pre-2021 data; the authors present three ways of fitting a generalised extreme value distribution and state that none of the three is fully satisfying [4]. The protocol’s own guardrails — restricting the shape parameter to roughly the interval between minus and plus 0.4, and using approximately the highest 10 to 20 per cent of data for threshold methods — are pragmatic devices for suppressing unphysical fits, not derivations from theory [3].

Model resolution is the second. AR6 states that explicit representation of severe convective storms requires non-hydrostatic models with horizontal grid spacings finer than four kilometres, and adds that even in convection-permitting models it remains difficult to simulate tornadoes, hail and lightning directly [1]. The 2026 National Academies report treats higher-resolution global models as potentially transformative for attribution, while listing limited model capability for small-scale regional events among the field’s persistent obstacles [8].

Rapid attribution and the price of being on time

Rapid attribution exists because the window in which an attribution result can inform recovery, insurance and policy is measured in days. The World Weather Attribution group reports in its own account of the practice that conventional studies take a year or longer while their process delivers findings within days or weeks, and states that its analyses are made public in that window with a commitment to submit to full peer review when the methodology is novel [12]. That is an assertion by the practitioners about their own process and should be read as such. The published protocol is the check on it: it is peer-reviewed, it fixes the eight steps in advance, and it makes the sequence auditable after the fact [3].

Two things in that account deserve weight. The group reports that defining the event proved both much harder and more important than initially expected, and that the attribution analysis itself is the easiest of the eight steps [12]. And it reports null results — its São Paulo drought analysis found the drought had not been made more severe by climate change, identifying vulnerability and exposure trends as the dominant drivers instead [12]. A rapid-response programme that only ever finds a signal would be evidence of a broken filter; one that publishes nulls is behaving like a measurement system.

A brass datum bolt set flush into the pale limestone sill of a tide house with a plumb bob hanging just above it, the bob still swinging in a small arc and its point not yet steady over the centre of the bolt
Figure 5. A result delivered in days is auditable only because the datum and the procedure were fixed and checked long before the event arrived.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

The 2026 National Academies committee does not endorse the trade unconditionally. It recommends that operational attribution programmes conduct periodic peer review to strengthen confidence in their results and to identify where capability is missing for understudied event types [8] — a recommendation that only makes sense if pre-publication review has in fact been traded away for speed.

What the field warns against

Van Oldenborgh and colleagues, writing as practitioners, catalogue the failure modes: event definitions chosen after the fact, confusion of trend detection with event attribution, inadequate or over-aggressive model evaluation, unacknowledged observational limitations, unstable statistical fitting choices, and communication that slides from a change in probability to a claim about causation [5]. AR6 adds selection bias at the level of the whole literature: the studied events are not a representative sample of events that occurred [1].

Three specific misreadings follow. An infinite risk ratio does not mean the event was impossible before; it means the counterfactual ensemble contained no such case, which is a statement about sample size as much as about physics [1, 11]. A fraction of attributable risk of 0.8 does not mean 80 per cent of the damage was anthropogenic; damage is set by exposure and vulnerability, which is why AR6 separates the difficulty of attributing a disaster from the tractability of attributing a hazard [1]. And an absence of attribution is not evidence of absence. Low confidence in tornado attribution reflects a short record and a four-kilometre grid, not a finding that convective hazards are unaffected [1, 7].

What would change this reading

Analysis. The confidence ordering across event classes is stable across the 2016 and 2026 National Academies assessments and AR6, and it tracks the directness of the thermodynamic link rather than the volume of literature. That is the strongest available evidence that it reflects physics rather than fashion.

Prediction, with a horizon. By 2031, I expect the confidence ordering to be unchanged — heat highest, convective hazards lowest — while attribution coverage in data-sparse regions improves through model-based analysis rather than through longer records, since records cannot be lengthened retroactively. This assumes continued growth of kilometre-scale global simulation and no collapse in observational network funding. Observable indicators would be attribution studies for African and South Asian events appearing in the peer-reviewed literature at rates approaching those for Europe, and hail or tornado attributions moving from low to medium confidence in a major assessment.

Disconfirmation. The prediction fails if a major assessment before 2031 raises confidence in tornado or hail attribution to medium or higher on the strength of convection-permitting simulation, which would show that resolution, not record length, was the binding constraint after all. It also fails if the storyline and probabilistic families converge on an agreed protocol for combining conditional and unconditional statements, since AR6 currently treats them as complementary rather than reconciled [1].

A tide gauge is the right instrument to keep in mind. Any single reading of the float is worthless on its own. What makes the reading mean something is the century of readings behind it, and the fact that the well was cut to a known depth and the drum turned at a known rate. Event attribution is the same trade. It converts a question about one afternoon into a question about a distribution, and it can only answer to the extent that the baseline was kept.