Ask a room of professional planners which method predicts the future best, and the honest answer is that the question is malformed. “The future” is not one target. A forecasting tournament question (“will X happen by date Y, yes or no”) and a corporate scenario-planning exercise (“what are the three or four internally consistent worlds we might be operating in a decade from now”) are not competing answers to the same problem. They are different instruments built for different jobs, and a third tradition — prediction markets, which convert dispersed private information into a single traded price — is different again. This article compares the three on dimensions where a real comparison is possible: how each has actually scored when tested, what kind of decision each output is built to support, how institutions have adopted each one, and how each behaves when the future does something structurally new rather than more of the same. It does not crown a winner, because the three are complementary tools with different domains of competence, not rival theories of the same phenomenon.

Fact, claim, analysis, scenario, prediction — kept separate

This article uses five registers and tries not to blur them. A fact is a documented result: a tournament score, a published calibration curve, a dated report. A vendor or institutional claim is a statement an organization makes about its own product or process — Shell’s account of what its scenarios are for, Good Judgment’s description of its own track record — reported as a claim, not independently re-derived here. Analysis is this article’s own reasoning connecting facts to a comparison. A scenario is an explicitly hypothetical, non-predictive “what if” used to stress-test decisions. A prediction is a specific, falsifiable claim about the future with a stated horizon, assumptions, and a condition that would prove it wrong. Readers should be able to tell at every point which register a sentence belongs to.

The three traditions, briefly

Probabilistic tournament forecasting asks individuals or teams to assign explicit numerical probabilities to well-defined, resolvable questions (“will country A hold elections before date B”), then scores those probabilities after the fact against what actually happened. The Good Judgment Project (GJP), led by Philip Tetlock and Barbara Mellers, was one of five university teams competing in IARPA’s Aggregative Contingent Estimation (ACE) program, a four-year geopolitical forecasting tournament run from 2011 to 2015 [3]. GJP’s approach — recruiting volunteer forecasters, training them briefly in probabilistic reasoning, teaming the best performers, and algorithmically aggregating and extremizing their estimates — was declared the tournament’s most accurate method, reportedly outperforming the other four university teams and, according to Good Judgment’s own retrospective, control groups that included intelligence-community analysts with access to classified information [4] [1]. That comparative claim against classified-access analysts comes from Good Judgment’s own published account of the tournament rather than from a fully independent third-party audit, and is reported here as an institutional claim rather than as an independently re-verified fact.

ADVERTISEMENT

Narrative scenario planning does not attempt to forecast a single outcome at all. It constructs a small number (usually two to four) of internally consistent, substantively different stories about how a set of driving forces outside an organization’s control might combine, and uses each story to pressure-test a strategy or a set of investment decisions. The tradition’s best-known corporate exponent is Shell, whose scenario planning group has published long-range scenario sets since the 1970s; its 2023 “Energy Transformation Scenarios” set out three named narrative pathways — Waves, Islands, and Sky 1.5 — differing in geopolitical alignment and the pace of decarbonization, explicitly not offered as a probability-weighted forecast of which pathway will occur [6]. The Delphi method, developed at RAND by Norman Dalkey and Olaf Helmer in the early 1950s to elicit and iteratively converge expert judgment on long-range military and technological questions, is an earlier and more structured cousin of scenario work: anonymous, iterated rounds of expert estimation with controlled feedback, aimed at consensus rather than at a scored probability [5].

Prediction markets ask neither an individual to state a probability nor a team to write a narrative. They let people trade contracts whose payoff depends on a future event, so that a contract’s price becomes a continuously updated, incentive-weighted estimate of the market’s aggregate belief. Platforms structured this way today include Manifold, Metaculus (a forecasting-aggregation platform rather than a real-money market in most of its questions), Kalshi, and Polymarket. Markets have also been used inside science itself: a 2015 PNAS study set up prediction markets on 44 psychology studies from the Reproducibility Project: Psychology and found the markets correctly anticipated replication outcomes and outperformed a parallel survey of individual researcher predictions on the same studies [7].

Dimension one: calibration under test conditions

Calibration is the property of a forecaster or system whose stated probabilities match observed frequencies — when it says “70 percent,” the thing happens roughly seven times in ten. It is measurable only where outcomes eventually resolve and scores can be computed, which is why tournament forecasting has the deepest audit trail of the three approaches: it was purpose-built to be scored, using rules like the Brier score, which compares a forecast’s stated probability to the realized outcome and rewards both accuracy and confidence calibrated to that accuracy. Mellers and colleagues, reporting on GJP’s internal tournament seasons, found that probability training, team-based collaboration among top forecasters, and performance tracking (feeding forecasters their own scores over time) each independently improved both calibration and resolution — the ability to distinguish likely from unlikely events with confidence rather than clustering near 50 percent [2]. A later synthesis in Current Directions in Psychological Science summarized the broader multi-team tournament literature the same way: aggregation algorithms, training, and teaming are the mechanisms behind the accuracy gains, not any single forecaster’s raw talent [1]. A third-party review by AI Impacts of the publicly released GJP data corroborated the general shape of these findings while noting that some specific magnitude claims in popular accounts of the tournament are harder to independently reproduce from the public data than the original papers suggest [10] — a useful caution against treating every widely repeated tournament statistic as equally solid.

Prediction markets are calibrated differently: by continuous price discovery rather than by discrete scored questions submitted to volunteers. Manifold’s own published calibration page shows its markets tracking close to the diagonal (stated probability matching realized frequency) across many thousands of resolved questions, which is a legitimate self-reported calibration record, not an independent audit [8]. The harder open question — separate from short-term calibration, which multiple platforms document reasonably well on well-specified, near-term questions — is whether that same calibration holds for long-horizon, low-liquidity, or definitionally fuzzy questions, where thin trading volume and ambiguous resolution criteria weaken the mechanism that makes market prices informative in the first place. That is an open empirical question this article treats as unresolved, not as settled in either direction.

Narrative scenarios are not calibrated at all, and are not supposed to be — they carry no probability to score. This is not a weakness relative to the other two approaches; it reflects a different job. A scenario set is judged by whether it expanded the range of strategies an institution considered and whether at least one of its scenarios later resembled events closely enough to have been useful preparation, which is an assessment of usefulness after the fact, not a Brier score computed at all.

ADVERTISEMENT

Dimension two: what kind of decision each output is built to support

A tournament-style probability (“62 percent this policy passes within eighteen months”) is built for decisions with a discrete trigger: fund the contingency plan or don’t, hedge the position or don’t, escalate a response or wait. It compresses uncertainty into a single number precisely so that a single go/no-go decision can hang on a threshold.

A scenario set is built for a different decision entirely: strategy selection under structural uncertainty about which world will obtain, when the organization cannot assign confident probabilities to the alternatives and does not want a single point estimate driving a large irreversible commitment. Shell’s own framing of its scenarios work is explicit that the exercise functions as a discipline against surprise and a stress test for strategies robust across more than one plausible future, rather than as a forecast to be graded later against the world that actually arrived [6]. That framing is Shell’s own characterization of its process, reported here as an institutional claim.

A market price is built for a third kind of decision: continuous position-sizing and information aggregation across many dispersed, self-interested participants, useful where an institution wants a live read on collective belief (and where that belief may shift daily) rather than either a single expert’s number or a slow-moving narrative document revised every few years.

None of these three outputs substitutes cleanly for either of the other two, which is the central reason a comparison should not try to declare an overall winner: a Brier score has no natural application to a scenario narrative, and a scenario matrix has no natural application to sizing an overnight trading position.

Dimension three: institutional adoption

Adoption differs by sector for reasons connected to the decisions above. National intelligence and defense analysis communities were the direct sponsors of the ACE tournament program and have since referenced structured forecasting methods, including elements of the “superforecasting” toolkit, in their own analytic-tradecraft training discussions, though the extent of formal operational adoption inside classified analytic workflows is not something this article can verify from open sources and is not claimed here. Corporate long-range planning, especially in capital-intensive, long-asset-life industries such as energy, has adopted narrative scenario planning far more visibly and for far longer: Shell’s scenarios practice runs in public, named cycles stretching back roughly five decades, most recently the 2023 Energy Transformation Scenarios [6]. Financial and forecasting-community adoption of market mechanisms has grown through platforms like Kalshi (a CFTC-regulated exchange) and Polymarket, alongside the explicitly non-monetary aggregation platform Metaculus, though public reporting on the 2024 U.S. election cycle noted meaningfully different accuracy outcomes across platforms with different liquidity, regulatory status, and trader composition — a reminder that “prediction markets” is not one mechanism but a family of them with different microstructure. Comparing platforms is worthwhile, and comparing them against each other is analysis this article does not attempt beyond noting that the family is heterogeneous.

Dimension four: discontinuities and structural breaks

All three approaches share a known weak point: they perform worse, by construction, on genuine discontinuities — events with no comparable historical base rate, sharp regime changes, or “unknown unknowns” that no one thought to write a question about. A tournament forecaster can only score a question that was posed; a market can only price a contract that was listed; a scenario set can only include the branches its authors thought to draw. This is not a defect unique to any one method — it is a structural limit of any approach built on extrapolation from observable indicators, historical base rates, or currently traded contracts. Scenario planning has some claim to a comparative advantage here, precisely because its explicit purpose is to imagine several structurally different worlds rather than to extrapolate a single most-likely path — but that advantage is only as good as the scenario authors’ imagination, and a narrative approach has no scoring mechanism to reveal, after the fact, whether its scenarios actually bracketed what happened or whether all of them missed high on the same side.

ADVERTISEMENT

A synthesis, not a ranking

Reading the record straight: probabilistic tournament forecasting has the strongest evidence base for calibration on well-defined, near-to-medium-term questions, because it was the only one of the three built from the outset to be scored that way, and its accuracy gains are attributable to specific, replicated mechanisms — training, teaming, and tracking — documented across multiple published tournament seasons [2] [1]. Narrative scenario planning has the strongest institutional track record for structural strategic decisions under deep uncertainty, precisely because it does not force a single probability onto questions where confident probabilities are not available, and its multi-decade adoption inside a capital-intensive industry is a real (if self-reported) institutional signal [6]. Prediction markets have the strongest claim to continuously updated, low-latency aggregation of dispersed belief, evidenced in a controlled scientific-replication setting [7] and in short-horizon calibration self-reports [8], but carry the least-resolved open question of the three regarding long-horizon and thinly traded questions. Treating any one of these as a general-purpose substitute for the other two is the error this article is built to avoid.

Falsifiable indicators for the next decade

Four observable indicators would let a reader judge, without waiting for another decade of hindsight, whether any one of these approaches is gaining or losing institutional ground.

First: if intelligence and policy institutions begin publishing (even redacted) internal Brier-score audits of their own analysts against tournament-trained forecasters, that would be observable evidence of tournament-forecasting adoption inside classified environments beyond what current open reporting can verify — the disconfirming condition is continued silence or explicit internal rejection of scored forecasting as a management practice over the next five years.

Second: if a major capital-intensive-industry scenario practice (energy, insurance, defense-procurement planning) publicly abandons named-scenario strategic planning in favor of single-point probabilistic roadmaps within the next decade, that would be evidence scenario planning’s comparative advantage under deep uncertainty is eroding — the disconfirming condition is continued or expanded publication of new named-scenario cycles by firms like Shell.

Third: if published platform-comparison studies (of the kind already appearing for the 2024 U.S. election cycle) begin consistently showing regulated, higher-liquidity markets calibrating better than lower-liquidity ones on matched questions across multiple election and policy cycles, that would support liquidity as the dominant driver of market forecast quality — the disconfirming condition is comparable accuracy regardless of liquidity once resolution criteria are held fixed.

Fourth: if any of the three traditions develops and publishes a scored method for auditing its own performance specifically on discontinuous, no-precedent events (rather than only on routine, well-precedented questions), that would be a genuine methodological advance worth tracking — the disconfirming condition is that no such audit method appears and all three traditions continue to be evaluated, as they are today, primarily on routine and precedented questions where extrapolation already works reasonably well.

A note on automation and the future of forecasting labor itself

One further comparison is worth making explicit, because it bears on how durable each tradition’s institutional role is likely to be: large language models are now used experimentally as forecasters in their own right, generating probability estimates on the same tournament-style questions GJP once used. This is squarely in scenario-and-prediction territory rather than settled fact, so it is treated that way here. The relevant economic frame is not “will software replace forecasters” but the more precise question Acemoglu and Restrepo pose for automation generally: does a new capability primarily displace an existing task, or does it also create new tasks that reinstate labor demand elsewhere in the same process [9]? Applied to forecasting, a plausible near-term scenario is that machine-generated probability estimates displace some of the routine, high-volume, well-specified questions that dominated early tournament work, while creating new tasks around auditing model calibration, designing resolution criteria robust to an automated respondent, and adjudicating disagreements between machine and human forecasts on genuinely ambiguous questions. A competing scenario is that automated forecasting simply commoditizes the tournament-forecasting tradition’s core product — a scored probability on a well-defined question — while leaving scenario planning and market mechanisms comparatively unaffected, because neither of those two traditions depends on an individual human’s probability estimate as its unit of production in the same way. This article does not adjudicate between the two scenarios; it flags the question as one where the next several tournament cycles, run with mixed human and machine participants, should produce observable calibration data within a few years rather than remaining purely speculative.

Reading a claim from any of the three traditions

A practical habit, drawn from the sourcing discipline above, transfers to any forecasting claim a reader encounters outside this article. Ask five questions in order. What exactly was scored — a discrete resolved question, a narrative that was never scored, or a price that moved continuously? Against what base rate or comparison group — an unstated intuition, a control group of professional analysts, or nothing at all? Verified by whom — the method’s own proponents publishing a retrospective, or an independent third party re-deriving the same number from public data? Over what horizon — days, months, or decades, since calibration evidence for one horizon does not automatically transfer to another? And finally, what decision was the output actually built to support — because a well-calibrated probability answers a different question than a well-constructed scenario, and neither answers the question a market price is built to answer. Applying that five-part check consistently is a better long-run habit than searching for a single “best” futures method, because the honest answer, defended throughout this article, is that the three traditions compared here are not competing for the same job.

What this comparison does not settle

This article does not resolve which approach an individual institution should adopt — that depends on the decision being supported, the time horizon, whether resolvable questions can be written at all, and how much the institution can tolerate a wrong point estimate versus a narrative that failed to include the world that actually happened. It does not establish that any one method is improving faster than the others; the evidence base for tournament forecasting is simply older and more thoroughly scored, which is a fact about publication history, not necessarily about which method best serves a given decision today. And it does not adjudicate the specific numerical superiority claims that circulate in popular accounts of the Good Judgment Project’s tournament results, several of which rest on the project’s own published retrospectives rather than on fully independent replication [4] [10]. Readers evaluating any forecasting claim — from any of the three traditions — should ask what was scored, against what base rate, verified by whom, and over what horizon, before treating a comparison like this one as more than a map of where the evidence currently stands.

A stack of printed Brier-score scorecards on clipboards, the top one half-annotated in red pen, beside a small response terminal mid-tally
Figure 1. Calibration is scored, not asserted: a Brier score compares a stated probability against what actually happened, aggregated across hundreds of questions.Image prompt and art direction by Brecht Corbeel; generation pending.
A large pinned scenario matrix on stiff board with two crossing axes, one quadrant card half-lifted from its clip as if being repositioned
Figure 2. Narrative scenario planning does not forecast a single outcome; it builds a small set of internally consistent, divergent worlds against driving forces an institution cannot control.Image prompt and art direction by Brecht Corbeel; generation pending.
Through a frosted-glass partition, a small server rack with status LEDs feeding a live market-price screen, one LED caught mid-blink
Figure 3. Prediction markets convert dispersed private belief into a single number through continuous trading rather than through a survey or a workshop.Image prompt and art direction by Brecht Corbeel; generation pending.
A row of small matte-plastic forecasting response terminals on a pale desk, one terminal's slider caught mid-drag between two probability values
Figure 4. Elicitation design shapes accuracy before aggregation ever begins: how a question is scored changes what a calibrated answer looks like.Image prompt and art direction by Brecht Corbeel; generation pending.
A stack of bound scenario-narrative booklets on a conference table, the top one open and caught mid-page-turn
Figure 5. A finished scenario set is a narrative document meant for a boardroom, not a probability distribution meant for a scoring rule — the two outputs are built to answer different questions.Image prompt and art direction by Brecht Corbeel; generation pending.