Before any technology forecast is worth reading, its method has to be named: base rates, scenarios, and markets each answer a different question, and each fails in a different way.

A forecasting-tournament scoring room mid-cycle: yesterday's questions tallied, today's still open. — Image prompt and art direction by Brecht Corbeel; generation pending.
This article opens a series on futures methods and technological scenarios by establishing, from scratch, what forecasting and scenario work actually are before any specific technology is discussed. It separates three distinct tools that are routinely conflated in popular writing about the future: reference-class (base-rate) forecasting as studied in Philip Tetlock's tournament research, scenario planning as practiced since Herman Kahn's work at RAND and Pierre Wack's work at Shell, and prediction markets as a mechanism for aggregating dispersed information into a price. Each method answers a different question, carries different failure modes, and produces a different kind of output — a probability, a set of divergent stories, or a market-clearing number — and none of them produces certainty. The article also sets out the discipline this series will use throughout: distinguishing verified fact from vendor claim, from analyst judgment, from constructed scenario, from testable prediction, so that later articles applying these tools to specific technologies can be read against a stated baseline instead of a vague sense that "experts expect."
Every claim about the technological future belongs to one of five categories, and almost no popular writing about the future bothers to say which. A statement can be a verified fact (“this chip shipped in this quantity in this year”), a vendor claim (“our roadmap projects X by 2030”), an analyst’s judgment (“we believe adoption will be slower than the roadmap implies”), a constructed scenario (“in one internally consistent future, adoption stalls for reason A; in another, it accelerates for reason B”), or a testable prediction (“there is a 70% chance of outcome X by date Y, conditional on Z, and here is what would prove this wrong”). This series opens with the discipline that keeps those five categories from collapsing into one undifferentiated mass of “experts say,” because the tools that produce futures claims — base-rate forecasting, scenario planning, and prediction markets — are three different instruments answering three different questions, and conflating them is the single most common error in technology forecasting.
This is a first-principles introduction: it assumes no prior exposure to forecasting research and builds the vocabulary this series will reuse in every subsequent piece, before any specific technology enters the discussion.
The most underused tool in forecasting is also the simplest: before estimating anything about a specific case, ask how similar cases have historically turned out. This is called reference-class forecasting, and it exists because of a documented, replicated bias in individual judgment. Daniel Kahneman and Amos Tversky identified what they called the planning fallacy — a systematic tendency for people estimating a specific project to focus on the details of that project (the “inside view”) while ignoring the distribution of outcomes for similar projects in the past (the “outside view”). Bent Flyvbjerg’s subsequent work on megaproject cost and schedule overruns operationalized this into a formal method: instead of asking “how long will this project take,” pull the completion times of a defined reference class of comparable past projects and use that empirical distribution as the anchor, adjusting only with strong justification [8].
Reference-class forecasting does not eliminate uncertainty; it replaces guesswork about a single case with an empirical distribution drawn from many cases. The hard part is defining the reference class honestly. A vendor forecasting its own product’s adoption curve, and an outside analyst using the adoption curves of the ten most comparable prior products, are performing two entirely different exercises even when both produce a single number. This is the first distinction this series insists on: a projection built from a stated reference class is not the same kind of claim as a projection built from an internal roadmap, and the two should never be quoted interchangeably.

Figure 1. Reference-class forecasting: a new estimate drawn from a drawer of comparable past outcomes, not from the plan alone. — Image prompt and art direction by Brecht Corbeel; generation pending.
The most direct empirical test of whether individual judgment can be systematically improved came from a set of government-run forecasting tournaments. The Intelligence Advanced Research Projects Activity’s Aggregative Contingent Estimation (ACE) program ran from 2010 to 2015 and set out to “dramatically enhance the accuracy, precision, and timeliness of intelligence forecasts” by testing methods for eliciting, weighting, and combining forecasts from many participants on real geopolitical questions with verifiable outcomes [4]. Several university teams competed; the team led by Philip Tetlock and Barbara Mellers, called the Good Judgment Project, consistently outperformed the others and, according to published results, beat a control group without any special training by around 60% and outperformed even intelligence-community analysts with access to classified information on the same questions [4].
The published research from that program identified three specific drivers of the improvement, not a single “superforecaster gene”: training that corrected known cognitive biases and taught forecasters to start from a reference class before adjusting; teaming, which let forecasters share information and challenge each other’s reasoning; and tracking, which selected top performers into elite teams based on demonstrated calibration over many questions [1]. The tournament format itself mattered as much as any individual skill: it forced every forecast into a comparable, scored, dated form, which is precisely what casual expert commentary about the future never does [2]. Tetlock and Dan Gardner’s later book Superforecasting popularized these findings, describing the recurring habits of the highest-scoring forecasters — breaking questions into components, updating in small increments as new information arrived, and treating a forecast as a number to be scored, not an opinion to be defended [3].

Figure 2. Tracking accuracy over time separates the merely confident forecaster from the calibrated one. — Image prompt and art direction by Brecht Corbeel; generation pending.
Analysis, not fact: it does not follow from this research that any individual reader can become a superforecaster by reading a list of habits; the tournament results describe a population-level, cross-question pattern measured over years of scored questions, and the research itself emphasizes that the gap between top and average forecasters was narrow on any single question and only became reliable when averaged across many [1].
A scenario is not a prediction. This distinction is the one most often lost when scenario planning is summarized for a general audience, so it is worth stating plainly: a scenario set does not assign probabilities to a menu of futures and is not trying to identify the most likely one. It is a structured way of building several internally consistent, sharply different futures so that a decision can be tested against all of them at once, revealing which choices are robust and which are fragile to a single assumption.
The method traces to two documented origins. Herman Kahn developed scenario-based analysis at the RAND Corporation starting in the late 1940s, applying systems analysis and game theory to think through alternative futures for nuclear strategy; Kahn is widely credited, alongside Pierre Wack, as a founding figure of the method [6]. Wack then carried the technique into the private sector at Royal Dutch Shell in the early 1970s, building scenarios that described divergent oil-market futures rather than a single forecast. According to accounts of that period, Shell’s planners had, by 1972, built out scenarios that included a sustained oil-price shock as one branch; when the 1973 oil crisis followed, the company was reportedly better prepared to respond than competitors who had planned around a single projected price [5].
Vendor and institutional framing to flag explicitly: both RAND’s and Shell’s own retrospective accounts of these episodes are institutional narratives told by the organizations that benefited from the story of their own foresight, and this series treats them as documented historical accounts of method and outcome, not as proof that scenario planning reliably prevents surprise in general. What scenario planning demonstrably does is force decision-makers to state the assumptions under which a given strategy fails, which is a different and more modest claim than “it predicts the future.”

Figure 3. Scenario planning does not predict one future; it builds several internally consistent ones to stress-test a decision. — Image prompt and art direction by Brecht Corbeel; generation pending.
A well-built scenario set has a small number of identifiable properties: a handful of scenarios (rarely more than four), built around two or three genuinely uncertain and high-impact driving forces rather than dozens of minor variables; internal consistency, meaning every detail in a given scenario has to be plausible given every other detail in that same scenario; and explicit difference, meaning the scenarios are built to diverge sharply rather than cluster around a consensus middle. The output is not a forecast to be graded against reality later — it is a stress test for a decision, and the scenarios that turn out closest to what happens are not “winners,” because that was never the exercise.

Figure 4. Old scenarios are audited, not discarded: a planning archive is also a record of which stories held up. — Image prompt and art direction by Brecht Corbeel; generation pending.
A third instrument works by an entirely different mechanism: letting people bet real stakes on an outcome and reading the resulting price as a forecast. The economic logic, formalized by Justin Wolfers and Eric Zitzewitz, is that a market price aggregates dispersed private information held by many participants into a single number more efficiently than asking any one expert, because participants who believe the price is wrong have a direct financial incentive to trade against it until it moves toward their information [7]. Wolfers and Zitzewitz’s review found that, across the markets they examined, prices tended to outperform polls, pundits, and other readily available benchmarks on the same questions, though they also documented known limits: thin trading produces noisy, easily distorted prices; markets need a large and diverse enough pool of participants with real information to aggregate; and a market price reflects what participants are willing to bet, which is not automatically the same thing as the objectively correct probability, particularly when a small number of well-funded traders can move a thin market [7].

Figure 5. A prediction market turns dispersed private information into a single visible number by letting people bet on it. — Image prompt and art direction by Brecht Corbeel; generation pending.
The distinction between a prediction market and a forecasting tournament matters for how each should be read. A tournament like ACE scores individual named forecasters over time against resolved outcomes, which is what let researchers isolate training, teaming, and tracking as specific, testable drivers of accuracy [2]. A market produces one number with no named author and no persistent scored track record attached to it; it is a mechanism for aggregation, not a mechanism for identifying which participants are reliably good at forecasting. Both are more disciplined than an unstructured expert opinion, and neither is a claim about certainty.
Every one of these instruments has a documented failure mode, and stating them plainly is part of using the methods honestly rather than as a rhetorical stamp of rigor.
Reference-class forecasting fails when the reference class is drawn dishonestly or too narrowly — when a forecaster picks the comparison set that supports a conclusion already wanted, or when a genuinely novel case has no adequate historical analog at all, such as a technology with no real precedent in kind or in speed of deployment. Base rates also fail at genuine discontinuities: by construction, a reference class built from history cannot anticipate a break from that history, which is exactly the situation many technology forecasts claim to be in — a claim that should itself be treated with suspicion, because “this time is different” is also the single most common justification for a bad forecast.
Scenario planning fails when the scenario set quietly narrows to variations on a single expected future rather than genuinely divergent branches, or when an organization treats one scenario as “the plan” instead of testing decisions across all of them — the exact discipline failure the method was built to prevent. It also produces no scored track record by design, so an organization cannot audit whether its scenario-planning process is actually improving its decisions over time, only whether individual scenarios happened to resemble what occurred.
Prediction markets fail when liquidity is too thin for the price to reflect real aggregated information, when the participant pool is not diverse or informed enough on the specific question, or when the question itself is poorly specified — an ambiguous resolution criterion breaks the entire mechanism, because traders cannot price what cannot be adjudicated [7]. And forecasting tournaments, for all their rigor, measure calibration on the kinds of short-to-medium-horizon, cleanly resolvable geopolitical questions the ACE program used; the same population-level gains have not been established for open-ended, decade-plus technology questions where resolution criteria are inherently harder to state in advance [1].
A reasonable objection at this point is to ask why any organization bothers with these methods at all, given that none of them promises to be right. The answer is that the alternative to a stated method is not a better forecast — it is an unstated one, made the same way but without the parts that let it be checked. Intelligence agencies, central banks, and large infrastructure funders did not adopt reference classes, scenario exercises, and aggregation mechanisms because these tools eliminate error; they adopted them because an institution that makes the same kind of decision repeatedly needs a process it can audit and improve across decisions, not just get right or wrong once. The ACE program itself existed for this reason: it was a five-year government investment in finding out, empirically, whether structured elicitation and aggregation methods actually beat unstructured analyst judgment on the same questions, rather than assuming they did [4]. Shell’s scenario practice persisted for decades past the 1973 episode that made it famous, not because every subsequent scenario set anticipated the next shock correctly, but because the planning discipline of stating divergent futures and testing strategy against all of them had organizational value independent of any single scenario’s accuracy [5].
This matters for how a reader should treat any institution’s forecasting claims going forward, including in later articles in this series. An organization’s use of a rigorous method is evidence that its process has been disciplined in a specific, checkable way — not evidence that its specific conclusion is correct. The two are separable, and this series will keep them separate: noting when a method is well-specified is a statement about process, and remains true even in the cases where the resulting forecast turns out wrong.
Every later article in this series will apply one or more of these three instruments to a specific technological question, and every one of them will state, explicitly, which instrument produced which claim. A forecast will carry a horizon, the assumptions it depends on, the observable indicators a reader could check before the horizon arrives, and a stated condition under which the forecast should be considered wrong — the same discipline the ACE tournaments imposed by scoring forecasters against resolved outcomes [2]. A scenario will be presented as one of several divergent, internally consistent branches, never as the likely future. A vendor roadmap will be labeled as a vendor’s claim about its own plans, not as an independent forecast. And where credentialed experts genuinely disagree about a technology’s trajectory, this series will describe the disagreement and its stakes rather than adjudicate a winner neither the data nor the methods can support.
None of this makes forecasting certain. What it does is make a forecast auditable: a reader who returns to a claim after the horizon has passed should be able to check, specifically, what was claimed, on what basis, and whether the stated disconfirmation condition was met. That is the entire difference between a forecasting method and a guess dressed in the vocabulary of one, and it is the standard the rest of this series is written to meet.
It helps to see the three methods applied to one concrete question rather than described in the abstract, so consider a single deliberately narrow example: will a given class of automated software agents handle a defined share of a specific back-office task inside a stated window of years. A reference-class forecaster would not start with the vendor’s roadmap. They would first ask what reference class of prior automation technologies this most resembles — perhaps optical character recognition, or robotic process automation, or an earlier generation of rule-based customer-service bots — and pull the empirical adoption curves and stated-versus-actual timeline gaps from that class, producing a distribution of plausible adoption speeds before ever looking at the specific product in question. That distribution becomes the anchor; only well-justified, explicitly stated reasons would move the estimate away from it, exactly the discipline Flyvbjerg’s method imposes on infrastructure-project estimates [8].
A scenario planner asked the same question would refuse to produce a single adoption number at all. They would instead identify the two or three genuinely uncertain forces that would most determine the outcome either way — verification cost for agent errors, the pace of institutional process redesign around the technology, and the emergence of new tasks that offset any labor displaced — and build a small set of sharply divergent, internally consistent stories: one where verification costs stay high and adoption stalls at the margins of low-stakes tasks, one where verification costs fall quickly and adoption reaches deep into higher-stakes workflows, and perhaps a third where the technology succeeds narrowly but new task categories more than absorb the effect on employment. None of these scenarios is “the forecast.” Each is a test: which institutional decisions look sound across all three, and which only look sound in the one scenario someone happens to prefer.
A prediction market asked the same question would not analyze the mechanism at all. It would post a specific, resolvable question — a defined share, a defined task category, a defined date — and let anyone with relevant private information take a position, aggregating whatever those participants collectively know or believe into a single tradeable price. That price would be informative to the extent the market was liquid and its participants informed, and uninformative, or actively misleading, if the question was ambiguous or the pool of traders thin [7]. None of the three outputs — a distribution, a set of branching stories, a market price — is a substitute for the other two, and a technology forecast that quietly slides between them without saying so is not more informative for combining them; it is simply harder to audit.
A forecast that cannot be shown wrong in principle is not a forecast, whatever confidence it is delivered with. The tournament research that most directly established this point scored forecasters using a strictly proper scoring rule — the Brier score, which penalizes both overconfidence and underconfidence against the outcome that actually occurred — over many resolved questions, which is what let researchers distinguish genuinely calibrated forecasters from merely confident or merely lucky ones [1]. Applying that discipline outside a formal tournament means every forecast in this series will carry four stated parts: a horizon (by when), the assumptions the estimate depends on, observable indicators a reader could track before the horizon arrives, and an explicit condition under which the forecast should be treated as disconfirmed. A claim missing any of those four parts is, by the definition this series uses, not yet a forecast — it is an impression wearing a forecast’s clothing, and the distinction is exactly what separates the scoring room this article opened with from the ordinary business of prediction by assertion.
Originally published at https://absolutedigitalpublishers.com/articles/futures-methods-and-technological-scenarios-a-first-principles-introduction.