Ask a room of experienced professionals to put a number on next year’s uncertain event and most will refuse, on the reasonable grounds that a number implies false precision. This guide takes the opposite position, developed over four decades of research into the same objection: refusing to quantify a belief does not make the belief more accurate, it just makes it unscoreable. A forecast that says “possible” or “likely” cannot be checked against what happened. A forecast that says “35 percent” can be, if you commit to checking it — and the checking is the entire discipline. This is a practitioner’s guide to three methods that make futures work falsifiable rather than performative: a calibrated probabilistic forecasting exercise scored with a proper scoring rule, a structured scenario-planning workshop with a documented facilitation sequence, and a prediction market built around one resolvable question. None of the three predicts the future. All three produce a record that can be graded after the fact, which is the property that separates a forecasting method from a rhetorical style.

Fact, claim, analysis, scenario, prediction

Five kinds of statement appear below and they are marked as such throughout. A fact is something independently documented and verifiable now — a published score, a dataset, a historical event. A vendor or platform claim is an assertion made by the organization that built or sells the method, presented as their claim rather than as settled fact. Analysis is this article’s own reasoning connecting facts to conclusions. A scenario is one internally consistent story about a possible future, explicitly not a prediction. A prediction carries a time horizon, the assumptions it depends on, the observable indicators that would confirm it is on track, and the specific condition that would disconfirm it. Readers should be able to tell at every point in this piece which of the five they are reading.

Part one: the calibrated forecasting tournament

Why base rates come first

The starting discipline in applied forecasting is not intuition, it is the base rate: how often has this general class of event occurred before, absent any specific reason to think this instance is different. Philip Tetlock’s two-decade study of expert political judgment is the foundational fact establishing why this matters — tracking roughly 82,000 individual probability judgments from 284 professional forecasters between 1984 and 2003, he found that the average expert’s forecasts were barely better than chance, and that experts who organized their thinking around one big overarching idea (Isaiah Berlin’s “hedgehogs”) did measurably worse than those who drew on many small, sometimes-contradictory sources of evidence (“foxes”) [2]. The hedgehog failure mode is specific and recognizable: a single causal story is treated as sufficient, disconfirming evidence gets explained away, and the base rate of “how often do dramatic forecasts like this one actually happen” never enters the picture. A base rate is a discipline against exactly that story-first reasoning — it forces the forecaster to ask what happened in the reference class before asking what feels true about this particular case.

ADVERTISEMENT

Proper scoring rules: why “confident and wrong” costs more than “uncertain and right”

A forecasting exercise is only informative if it is scored with a rule that rewards honesty. Tilmann Gneiting and Adrian Raftery’s formal treatment of proper scoring rules is the mathematical fact underlying every tournament described in this guide: a scoring rule is “proper” if a forecaster’s expected score is maximized by reporting their true probability estimate, and “strictly proper” if that maximum is unique [3]. Two proper scoring rules dominate practice.

The Brier score is the mean squared error between a stated probability and the binary outcome (1 if it happened, 0 if it did not):

B=1Ni=1N(pioi)2B = \frac{1}{N}\sum_{i=1}^{N} (p_i - o_i)^2

where pip_i is the forecaster’s stated probability on question ii and oi{0,1}o_i \in \{0,1\} is the realized outcome. Lower is better; a perfectly calibrated coin-flip forecaster scores 0.25 on a 50/50 question, and a forecaster who says 100 percent and is wrong scores the maximum possible 1.0.

The logarithmic score, used by Metaculus, is log(p)\log(p) if the event occurs and log(1p)\log(1-p) if it does not. This is a platform claim and design note, not independent research: Metaculus documents that the log score is far more punitive of confident misses than the Brier score is rewarding of confident hits — moving from 99 percent to 99.9 percent gains almost nothing if correct but costs roughly 2.3 points if wrong [7]. That asymmetry is deliberate: it is what makes the rule “strictly proper” in Gneiting and Raftery’s sense, because it punishes the temptation to round an honest 90 percent up to a persuasive-sounding 99 percent.

Running the exercise

A calibrated forecasting exercise, as run inside the tournaments described below, has a repeatable structure:

ADVERTISEMENT
  1. Write questions with an unambiguous resolution criterion and date before any forecast is made. “Will X happen by date Y, as determined by source Z” — not “will X happen,” which invites post-hoc argument about what counted.
  2. Elicit a probability, not a verbal hedge. Forecasters commit to a number, typically to the nearest percentage point, which is what makes scoring possible at all.
  3. Update on a visible timeline, not only once. Recording when a forecaster moved their estimate, and on what evidence, is itself scoreable — Barbara Mellers and colleagues found that top performers updated more frequently and in smaller increments than average forecasters, rather than making one confident call and holding it [1].
  4. Score every resolved question with a fixed, pre-declared rule — Brier or log score — and never after the fact substitute a rule that would have made a favored forecaster look better.
  5. Aggregate across many forecasters, weighting toward historically accurate ones, rather than trusting any single expert’s number. IARPA’s Aggregative Contingent Estimation program, which ran from 2011 to 2015 and generated the Good Judgment Project data, existed specifically to test aggregation and elicitation methods against each other under this kind of scoring discipline [4].

What the tournament data actually showed

This is the fact layer, and it is worth being precise about what was and was not found. Mellers and colleagues, reporting on the IARPA-funded Good Judgment Project, identified a subset of forecasters — roughly the top two percent — whose accuracy was not a one-season fluke: when regrouped into elite teams and given continued access to feedback, they maintained higher accuracy than a randomly selected regression-to-the-mean pattern would predict [1]. IARPA’s own program materials describe the broader four-year competition among elicitation and aggregation methods, and its subsequent public release of the underlying forecast-level data through Harvard’s Dataverse allows outside researchers to re-score the tournament under different rules — a transparency step that most forecasting claims in industry and government do not offer [5]. Good Judgment’s own public platform, Good Judgment Open, continues to run this tournament format commercially and for research sponsors, badging its top performers as “Superforecasters” — that badge and the marketing language around it is a vendor claim, not an independently audited credential, even though the underlying method it is built on is peer-reviewed [6].

Individual calibration training

Before a person’s aggregate forecasts mean anything, their individual calibration needs to be checked: of all the times they said “70 percent,” did the event happen roughly 70 percent of the time? A calibration curve built from many such judgments, plotting stated probability against realized frequency, is the tool; a forecaster whose curve sits above the diagonal is overconfident, below it is underconfident. Training runs typically use trivia questions with fast resolution — the mechanism generalizes because it is scoring the belief-reporting habit, not domain knowledge.

A single forecaster's desk with a probability-wheel calibration trainer, an index card, and a pen resting mid-entry
Figure 1. A calibration trainer at a single desk: the wheel gives an outcome, the forecaster records the probability they actually believed, not the one they hoped for.Image prompt and art direction by Brecht Corbeel; generation pending.

Practically, this means treating every stated forecast as a data point to be graded later, not a one-time verdict to be delivered and forgotten.

A corkboard of dated question cards with one card being unpinned and moved toward a resolved tray
Figure 2. Resolution day: a question card is unpinned from its due-date column and moved to the resolved tray, the outcome now fixed and the forecast now scored.Image prompt and art direction by Brecht Corbeel; generation pending.

Part two: the structured scenario-planning workshop

Where a forecasting tournament produces a probability, scenario planning deliberately does not. It produces several internally consistent stories about how a specific decision’s environment could unfold, explicitly refusing to assign likelihoods to them, on the theory that forcing a single most-likely story back into the process reintroduces the hedgehog failure the method exists to avoid. This is a documented facilitation method with a specific origin, not a loose brainstorming format.

The GBN eight-step method

Peter Schwartz, drawing on scenario work developed at Royal Dutch Shell and later at Global Business Network (GBN), documented an eight-step sequence in The Art of the Long View, describing exercises run for organizations including the White House, the EPA, and several corporations [8]. The sequence, as a fact about the documented method:

  1. Identify the focal issue or decision. Not “the future of the industry” in general, but the specific decision the organization actually has to make and by when.
  2. List the key local forces that would influence the success or failure of that decision — customers, competitors, regulators, suppliers.
  3. List the driving forces in the macro-environment — social, economic, political, environmental, technological trends acting on the local forces.
  4. Rank forces by importance and by uncertainty. This step is the hinge of the whole method: most forces are either fairly predictable or fairly unimportant, and the workshop’s job is to find the small number that are both important and genuinely uncertain.
  5. Select the scenario logic. The two highest-ranked uncertain-and-important forces typically become two axes, producing a 2x2 matrix of four scenario quadrants, though Schwartz’s method allows other logics when a matrix would force an artificial symmetry.
  6. Flesh out the scenarios into short narratives — not forecasts, stories — describing how the world got from today to each quadrant.
  7. Identify the implications of each scenario for the focal decision from step one.
  8. Select leading indicators and signposts — observable, checkable signals that would tell the organization which scenario is becoming reality as events unfold.
A large worktable with sticky notes being arranged onto two crossed axis strips, one note held mid-placement over a quadrant
Figure 3. A scenario workshop mid-session: driving-force notes sorted onto two crossed axes, one note still hovering over its quadrant, not yet pressed down.Image prompt and art direction by Brecht Corbeel; generation pending.

Where scenario planning earns its keep, and where it doesn’t

Analysis: the method’s value is not in guessing correctly which quadrant occurs — the explicit refusal to rank likelihoods means it is not graded that way, and there is no equivalent to a Brier score for a scenario set. Its value is in step 8: an organization that has pre-committed to naming which signposts distinguish its four futures can recognize a regime shift faster than one encountering the same evidence with no prepared framework, because the interpretive work was done in advance rather than under the time pressure of the event itself. That is a real and testable claim, but it is a claim about decision speed and preparedness, not about forecast accuracy, and conflating the two is the most common misuse of scenario output — treating one quadrant as “the prediction” defeats the method’s own design.

ADVERTISEMENT

The failure mode on the other side is treating the matrix as decoration: a scenario exercise that skips steps 1 and 4 — a real focal decision and a genuine uncertainty ranking — degrades into four plausible-sounding stories with no connection to what the organization actually has to decide, and no signposts anyone will actually watch for. The steps that get skipped under time pressure are consistently 4 and 8: ranking is uncomfortable because it means declaring some pet issues unimportant, and signposts are uncomfortable because they are a commitment to being provably wrong later.

Part three: constructing and reading a prediction market

A prediction market prices a claim the way a commodity market prices wheat: a contract that pays a fixed amount if an event happens and nothing if it doesn’t will trade at a price that reflects the balance of participants’ beliefs, weighted by how much they are willing to risk on being right. The method’s appeal is that it aggregates dispersed information through incentives rather than through a facilitator’s questions.

A cautionary fact before the method

Before describing how to build one, the clearest documented case of a prediction market failing for institutional rather than technical reasons: DARPA’s Policy Analysis Market, a 2003 proposal to let participants trade contracts on political and economic developments in the Middle East, was publicly announced and cancelled within roughly one day after senators characterized it as a “terrorism futures market” that could be read as profiting from attacks [9]. The underlying pricing mechanism was not shown to be technically flawed — the program was killed by the reputational and ethical objection to the subject matter, not by evidence the market itself mispriced anything. This is a necessary caveat for any institutional user: a technically sound prediction market on a sufficiently sensitive question can be shut down by the same actors it might have informed, and the method’s designers must screen questions for this failure mode as carefully as they screen for resolvability.

Constructing a market for a real question

  1. Write a resolution criterion identical in rigor to a tournament question — an ambiguous question produces an unreliable price regardless of the market mechanism underneath it.
  2. Choose a market structure. A continuous double auction (buyers and sellers post limit orders that match, as on Polymarket or Manifold Markets) versus a market maker that always quotes a price and adjusts it after each trade (used by Metaculus’s community-prediction aggregator and many academic implementations) trade off liquidity for simplicity; a market maker guarantees a tradeable price even with few participants, at the cost of the operator absorbing some risk.
  3. Decide the stake. Real-money markets (Polymarket, historically the Iowa Electronic Markets) and play-money markets (Manifold Markets) are both live implementations; this is a genuinely contested design question addressed below.
  4. Publish the current price continuously, not only at resolution — the informational value of a market is in its running price, which updates faster than a survey or a scheduled forecasting round can.
  5. Resolve on the pre-declared criterion and pay out accordingly, with the resolution source and evidence documented publicly so the market’s credibility compounds across questions rather than resetting each time.
A paper-tape market ticker and order-book printer with a curl of tape mid-feed and a price wheel just turned
Figure 4. A prediction-market ticker: the price wheel has just turned one notch as a new order clears, and the paper tape is still feeding out beneath it.Image prompt and art direction by Brecht Corbeel; generation pending.

Does the stake matter? Real money versus play money

This is a genuine empirical disagreement rather than a settled fact, and it should be characterized as one. The vendor-side claim from Manifold Markets is that its play-money platform performed close to real-money benchmarks: Manifold reports its 2022 U.S. midterm election forecasts were close in accuracy to FiveThirtyEight’s model and comparable to real-money markets on overlapping questions, and it publishes an ongoing site-wide calibration page showing stated probabilities against realized frequencies [10]. Independent commentary on this comparison notes that neither market type has been shown to be systematically more accurate across a broad sample of matched questions, but that the gap between play-money and real-money pricing does show up at the margins — specifically in thinly traded, low-volume questions, where real cash appears to draw in more careful traders willing to correct a mispriced contract, while a lightly traded play-money market can sit at a stale price longer. The honest summary: for well-traded, broadly followed questions, stake type appears to matter less than trader population size and diversity; for niche or low-attention questions, it likely still matters, and no side of this debate has produced a large matched-pair study settling it definitively.

A mechanical aggregation tray with many small weighted paper slips feeding down a funnel toward a single reading dial
Figure 5. Aggregation: many individual forecast slips, weighted by track record, feed down a funnel toward one combined reading, the last slip still in the chute.Image prompt and art direction by Brecht Corbeel; generation pending.

Reading a market price correctly

Analysis: a market price is not “the probability of the event” in any deeper metaphysical sense — it is the price at which the marginal trader is indifferent between buying and selling, which converges toward a good probability estimate only when the market has enough diverse, informed participants trading enough volume to correct mispricing. A thinly traded market on an obscure question can sit at a stale or manipulated price for a long time with nobody profitable enough to correct it; a heavily traded market on a widely followed question is much harder to move away from the crowd’s honest aggregate belief, because anyone who thinks the price is wrong has a direct financial incentive to trade it back. Reading a market price without checking its volume and participant count is the prediction-market equivalent of reading a single expert’s confident forecast without checking their track record.

Model error, discontinuities, and institutional use

None of these three methods handles a genuine discontinuity well, and this limitation should be stated plainly rather than papered over. Base rates, calibration training, scenario axes, and market prices are all built from a reference class of things that have happened before or forces already visible today. A forecasting tournament will systematically underweight a low-base-rate structural break precisely because it has no prior instances to calibrate against; a scenario workshop can only include a discontinuity if someone in the room thought to name it as a driving force in step 3; a market can only price a contingency that someone bothered to write a resolvable contract for. This is not a flaw unique to these methods — it is a general property of any method built from historical or currently visible information — but treating any of the three as a discontinuity detector is a category error. Their proper institutional use is to make the ordinary, extrapolable part of an uncertain future rigorous and checkable, freeing genuine judgment and dedicated red-teaming for the part that is not.

A prediction, stated explicitly

Prediction, with horizon and disconfirmation condition stated per the editorial standard above: over the next five years (horizon: 2026–2031), we expect institutional adoption of scored forecasting tournaments and constructed prediction markets to expand within organizations that already have some quantitative risk culture (finance, insurance, some public-health and defense agencies), and to remain rare inside organizations without one, because the methods’ central cost is not technical but cultural — they require an organization to publicly commit to being provably wrong. Assumptions: no major regulatory restriction on internal-use prediction markets comparable to the 2003 Policy Analysis Market reaction, and continued availability of platforms like Metaculus and Manifold Markets as reference implementations. Observable indicators consistent with this prediction: a rising count of corporate and government-sponsored tournaments on Good Judgment Open and Metaculus year over year; continued publication of resolved-question track records rather than only marketing claims. Disconfirmation condition: if scored tournaments and constructed markets fail to expand beyond the current small set of finance, insurance, and national-security-adjacent institutions even where quantitative risk culture already exists, or if a comparable public backlash to the Policy Analysis Market episode causes existing platforms to withdraw from institutional use, the prediction should be considered falsified.

Where the disagreement actually lies

Where practitioners genuinely disagree, that disagreement should be characterized rather than resolved by authorial fiat. Forecasting-tournament researchers and scenario-planning practitioners disagree, sometimes sharply, about whether the two methods are complements or substitutes: the tournament camp argues that a probability with a scored track record is strictly more useful than an unranked scenario set for any decision that can be reduced to a checkable question, while the scenario camp argues that most decisions that matter cannot be reduced to a single checkable question without losing the structural forces that made the decision hard in the first place, and that scoring accuracy on the wrong question is worse than admitting the question can’t be scored. Neither side has produced a controlled comparison settling which approach produces better organizational decisions, because “better decisions” is itself not agreed upon as a scoreable outcome. This guide does not adjudicate that dispute; it treats the two methods as answering different questions and recommends practitioners use base rates and scored tournaments where a question is genuinely resolvable, and a Schwartz-style scenario workshop where the decision’s environment is too structurally uncertain to reduce to one.

Summary for practitioners

Run a forecasting tournament when you have — or can write — a resolvable question, and you are willing to score every entrant on a pre-declared proper scoring rule rather than after the fact. Run a scenario workshop when the decision’s environment has two or more forces that are both highly uncertain and highly important, and you are willing to name leading indicators in advance rather than only after a future has already arrived. Build or read a prediction market when a question is resolvable, tradeable interest exists, and you have checked the market’s volume and participant diversity before trusting its price — and screen the question itself for the kind of reputational risk that killed the Policy Analysis Market regardless of how sound the pricing mechanism underneath it would have been. All three methods share the same underlying discipline: a number, a story, or a price is only as useful as the record kept of whether it was right, and that record has to be built before the outcome is known, not assembled afterward to flatter whichever method looked best in hindsight.