How forecasting became a discipline: RAND's Delphi rounds, Shell's scenario rooms under Pierre Wack, and the tournament that measured what actually predicts the future.

RAND's Delphi method, first published in 1963, circulated anonymous expert estimates through successive rounds rather than a single meeting [@rand-delphi-1963]. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.
Futures methods did not emerge from a single insight but from three distinct institutional programs, each responding to a different failure of ordinary forecasting. RAND Corporation's Delphi method, developed from classified Air Force work in the late 1940s and published openly in 1963, tried to extract a defensible consensus from expert opinion without letting status distort the answer. Shell's scenario planning practice, built by Pierre Wack's team in London through the early 1970s, abandoned point forecasts altogether in favor of structured alternative futures, and Shell's own account credits it with preparing management for the 1973 oil shock. The Good Judgment Project, run under IARPA's ACE tournament from 2011, did something neither RAND nor Shell had done: it scored forecasters against realized outcomes at scale and identified a measurable "superforecaster" skill. This article traces those three programs as verifiable history, separates documented fact from vendor claim and interpretation, and closes with conditional, falsifiable expectations for where institutional forecasting is headed next.
Ask a forecaster which of the following actually predicts the future — a room of hand-picked experts, a set of named alternative worlds, or a crowd of amateurs scored against reality — and the honest answer is that all three have been tried at institutional scale, and only one of them has ever been measured well enough to say how good it is. That measurement did not exist until 2011. Everything before it was method without a scoreboard.
This is a history of three specific, verifiable programs, not a general theory of prediction. RAND Corporation’s Delphi method tried to extract a defensible consensus from expert opinion under anonymity. Shell’s scenario-planning practice, built under Pierre Wack, abandoned point forecasts in favor of structured alternative futures. The IARPA-funded Good Judgment Project did something the other two never did: it tracked hundreds of thousands of individual probability estimates against outcomes that actually happened, and it identified, by name, a class of forecasters who were reliably better than everyone else. Read together, the three programs are less a lineage than three different answers to the same problem — how to make a defensible claim about an unknown future — and the record of what worked, what didn’t, and what was never actually tested is worth separating carefully.
RAND Corporation was incorporated as an independent nonprofit in May 1948, spun out of the Douglas Aircraft Company’s wartime “Project RAND” to connect military planning with research and development decisions [3]. Through the 1950s the organization became known for systems analysis — the discipline of linking data, formal models, and decision-making into one analytic pipeline — and it was inside that culture that the Delphi method was developed [3].
The method itself, as formally published by Norman Dalkey and Olaf Helmer in the RAND research memorandum “An Experimental Application of the Delphi Method to the Use of Experts,” addressed a specific, narrow problem: how to get “the most reliable consensus of opinion of a group of experts” without letting the discussion collapse into groupthink or status hierarchy [1]. The technique, in its original form, was procedural rather than clever. Experts answered a questionnaire individually and anonymously. Their answers, and the reasoning behind outlying answers, were summarized and circulated back to the same group. The group answered again, in light of the summarized spread of opinion, still without knowing who had said what. The process repeated for several rounds, and the estimates typically converged.
That convergence was, and remains, the method’s central ambiguity. RAND’s own retrospective commentary on Delphi, published decades later, frames it as a way of “generating evidence” in exactly the cases where hard data does not exist — technology timelines, health-policy judgments, disaster preparedness assessments — precisely because the alternative is either a single expert’s opinion or no structured elicitation at all [2]. That is a fair description of what Delphi does: it manufactures a documented, traceable consensus number where none existed. It is a materially different claim from saying that number is accurate. The original 1963 paper was an experimental demonstration of the elicitation procedure, not a validation of forecasting accuracy against later-realized outcomes [1]. Nothing in RAND’s own record from this period claims that Delphi-derived consensus estimates were later checked against what actually happened at any systematic scale — that check, when it eventually came, arrived from a different institution and six decades later. Treat that gap honestly: Delphi is a method for producing a defensible, procedurally fair group estimate; whether such estimates are well-calibrated forecasts is a separate empirical question the original program did not answer.

Figure 1. RAND separated from Douglas Aircraft in May 1948 to become an independent nonprofit research organization built around systems analysis [@rand-history-decade]. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.
RAND’s futures-methods work did not stop at Delphi. Herman Kahn, who worked at RAND through the 1950s before founding the Hudson Institute in 1961, extended a different technique from the same institutional culture: the named alternative future, or scenario, first developed publicly in his 1960 book On Thermonuclear War [7]. Kahn’s scenarios were not attempts to predict a single outcome. They were structured, internally consistent narratives meant to make military and civilian planners rehearse decisions under conditions they had not yet had to face — an approach RAND’s own later methodological literature still classifies as one of its core futures-analysis tools, alongside Delphi and other forecasting techniques [11]. Kahn’s scenario technique is the direct ancestor of the corporate practice that Shell would build a decade later, and the distinction between the two RAND-derived lineages — elicited consensus versus structured alternative narrative — is the first fork this history has to keep straight, because later practitioners frequently blur it.

Figure 5. Herman Kahn's scenario technique, developed at RAND and extended at the Hudson Institute from 1961, treated named alternative futures as a strategic-planning tool rather than a single prediction [@kahn-thermonuclear-wikipedia] [@rand-futures-methodologies]. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.
Shell’s scenario-planning group, built through the late 1960s and early 1970s under planner Pierre Wack, took the RAND-derived scenario idea and applied it to corporate strategy rather than nuclear deterrence. The premise, as Wack later described it in his two 1985 Harvard Business Review articles, was that a single-point forecast of oil prices, demand, or supply was not merely difficult — it was the wrong kind of object to produce, because it hid the assumptions that would make it wrong [4]. Instead, Wack’s team built a small number of internally consistent, structurally different futures and asked Shell’s management to reason through the strategic implications of each one, well before knowing which (if any) would occur.
Shell’s own retrospective history, “40 Years of Shell Scenarios,” credits this practice with preparing the company’s management for the October 1973 oil crisis — not by predicting the embargo’s timing, but by having already walked through a future in which oil-producing states used price and supply as a lever, so that when the crisis arrived, Shell’s decision-makers recognized the situation rather than treating it as an unprecedented shock [6]. Wack himself, in the Harvard Business Review account, is explicit that the value was in the “prepared mind” of the manager rather than in forecasting timing or magnitude correctly [4]. It is worth being precise about what kind of claim this is: it is Shell’s own institutional narrative about its own practice, corroborated by a named practitioner’s own published account, not an independent audit by a third party comparing Shell’s pre-1973 planning documents against its post-1973 competitive performance. That distinction matters, because “scenario planning helped us respond faster” is a claim a company is well positioned to make honestly and also well positioned to overstate; nothing in the public record cited here constitutes an independent, outcome-blinded test of the counterfactual (how Shell would have performed without the scenario exercises).

Figure 2. Pierre Wack's planning team in London built structured alternative futures rather than single-point forecasts, a practice Shell credits with preparing managers for the 1973 oil shock [@wack-hbr-1985] [@shell-40-years-scenarios]. — Image prompt and art direction by Brecht Corbeel; generation pending.
What scenario planning explicitly does not claim, and this is the second fork to keep straight, is calibrated probability. A scenario set does not assign a number to “how likely is this future”; it deliberately displaces attention from likelihood onto structural plausibility and strategic robustness across multiple futures at once. That makes scenario planning close to unfalsifiable as a forecasting method in the narrow sense — there is no scorable prediction to check later, only a strategic posture that either did or did not turn out to matter. This is not a flaw so much as a different tool built for a different job: a decision-support exercise, not a probability estimate. Confusing the two — treating a named scenario as though it carried an implicit likelihood, or worse, treating “we considered this scenario” as evidence the organization forecast the event — is a common misreading of Shell’s own account, and one this article deliberately avoids.
Between RAND’s Delphi work and the tournament era, a third body of research quietly undercut the assumption both Delphi and scenario planning shared implicitly: that domain experts, given a structured process, would be reasonably calibrated about the future. Philip Tetlock’s long-running study of professional political and economic forecasters — begun in the 1980s and published as Expert Political Judgment in 2005 — collected tens of thousands of predictions from people who made their living commenting on political and economic trends, and scored them against what actually happened [9]. The finding that survived into popular reference, that the average expert’s accuracy was statistically indistinguishable from chance, is widely summarized as “no better than a dart-throwing chimpanzee” [9]. More consequential than the headline number was a secondary finding: the experts’ confidence in their own predictions was inversely, not positively, correlated with accuracy, and pundits with the most media visibility and the most definite public style performed the worst [9].
That finding reframes both RAND’s and Shell’s programs retroactively. Delphi’s anonymity and iteration were designed to counter status effects and public commitment bias inside a panel — exactly the mechanism Tetlock later found undermining public forecasters outside any structured process. Tetlock’s study did not test Delphi panels directly, and it would be an overreach to claim it validates Delphi’s design; but it does supply independent evidence for the specific failure mode Delphi’s procedure was built to suppress, which is a meaningfully different and more defensible claim than treating the two studies as directly comparable.
In 2011, the US intelligence community’s Intelligence Advanced Research Projects Activity (IARPA) launched the Aggregative Contingent Estimation program, a multi-year, multi-team tournament explicitly built to compare forecasting methods against each other using real geopolitical questions with knowable, dated resolutions [8]. This is the detail that separates the ACE tournament from everything described above: for the first time, competing forecasting methods were run in parallel against the same questions, on the same time frame, and scored against outcomes that actually occurred, rather than against internal plausibility or client satisfaction.
The Good Judgment Project, led by Philip Tetlock and Barbara Mellers at the University of Pennsylvania, was one of several funded teams and became the tournament’s dominant performer across its four-year run, collecting on the order of a million individual forecasts across roughly 500 questions [8]. Good Judgment’s aggregated median forecasts beat the tournament’s control group by more than fifty percent and beat competing research teams by tens of percent, and its results were reported as exceeding even intelligence analysts working with access to classified information on the same questions [8]. From an initial pool of roughly 5,000 volunteer forecasters, the project identified around 260 individuals — later termed “superforecasters,” the top few percent by tracked accuracy — whose performance improved further once they were grouped into dedicated teams in the tournament’s second year [8].

Figure 3. Philip Tetlock's long-run study of expert political forecasts found accuracy no better than chance for many pundits, motivating later work on measurable forecasting skill [@tetlock-epj-braverangels]. — Image prompt and art direction by Brecht Corbeel; generation pending.
Independent summaries of the project’s internal working papers, compiled by researchers unaffiliated with Good Judgment itself, describe specific practices correlated with the group’s accuracy: frequent small updates to existing estimates rather than large infrequent ones, active aggregation of many individual forecasts rather than reliance on any single expert, and selection of forecasters based on tracked track record rather than credentials or seniority [10]. This is the closest thing to a controlled comparison this history has: a program that ran expert intuition, statistical aggregation, and tracked individual skill against one another on identical questions with a fixed resolution date, and reported which one won. It is still one tournament run by one funder over one four-year window covering primarily geopolitical and economic questions — a real result, not a universal law of forecasting, and its conclusions should be read as bounded by that scope rather than extended by assumption to domains (long-horizon technological change, for instance) the tournament did not test directly.

Figure 4. The IARPA-funded Good Judgment Project scored hundreds of thousands of forecasts against realized outcomes from 2011, identifying a top tier of forecasters later called superforecasters [@goodjudgment-track-record] [@aiimpacts-gjp-evidence]. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.
Laid side by side, the three programs answer different questions and license different conclusions.
RAND’s Delphi work established that a structured, anonymous, iterative elicitation procedure produces a different — typically more convergent — group estimate than an open discussion or a single expert’s opinion [1] [2]. It did not establish that the resulting estimate is well-calibrated against later outcomes at any systematic scale; that validation question was simply not the one the original research program was built to answer.
Shell’s scenario practice under Wack established, by the company’s own credible institutional account, that rehearsing multiple structurally different futures changed how quickly management recognized and responded to an unanticipated shock [4] [6]. It did not, and by design cannot, produce a scorable probability forecast, and the claim of causal benefit rests on the organization’s own retrospective narrative rather than an independent counterfactual audit.
Tetlock’s base-rate research established, with a large sample scored against real outcomes over two decades, that domain expertise and media prominence in political forecasting were poor and sometimes negatively correlated with accuracy [9]. The Good Judgment Project then established, in a purpose-built tournament with dated resolutions, that a specific combination of frequent updating, statistical aggregation, and track-record-based selection substantially outperformed both a control group and rival forecasting teams on geopolitical questions over a four-year window [8] [10].
One mechanism sits adjacent to all three programs without belonging cleanly to any of them: the prediction market, in which forecasters trade contracts that pay out based on a real-world outcome rather than submitting a probability directly. The theoretical appeal is that a market price aggregates dispersed private information through an incentive to profit from correcting a mispriced estimate, rather than through a facilitator’s questionnaire rounds or a tournament’s scoring rule. The Good Judgment Project’s own tournament design deliberately tested this question inside the ACE program: prediction-market-style mechanisms and prediction-polling mechanisms were run as separate conditions alongside the aggregated-forecaster teams, allowing IARPA to compare methods of combining judgment, not only individual forecasters, on the same underlying questions [8]. The reported result favored Good Judgment’s own combination of selected forecasters and statistical aggregation over the competing designs fielded in the same tournament window, which is a narrower and more specific claim than “markets don’t work” — it is a claim about which aggregation mechanism won a particular multi-year government-run competition, not a general verdict on market-based forecasting across all domains, and it should not be extended past that scope.
This matters for reading the three founding programs correctly, because the temptation, in each case, is to promote a documented institutional result into a universal claim about “how to predict the future.” RAND’s Delphi memoranda documented a procedure for group elicitation, not a proof that groups elicited this way beat markets, tournaments, or plain trend extrapolation on accuracy. Shell’s account documented an organizational benefit from rehearsal, not a quantified forecasting edge over its competitors. The ACE tournament documented that one specific combination of selection and aggregation beat several rivals on one class of questions over one funding cycle. Each claim is real and worth taking seriously. None of them licenses the others’ conclusions by association, and treating “futures methods” as a single unified discipline with one winning technique is the recurring misreading this history is built to head off.
A separate and underdiscussed thread runs through all three programs: institutional forecasting failures are rarely failures of arithmetic. They are failures of framing — asking the panel, the scenario team, or the tournament forecasters the wrong question, or defining the resolution criterion so loosely that a correct underlying judgment cannot be scored as correct. Tetlock’s base-rate research is instructive here precisely because it separates the two failure modes: the pundits in his sample were not uniformly bad at reasoning about evidence, but many were systematically overconfident about the scope of what their expertise let them claim, extending a valid regional or sectoral judgment into a confident prediction about a domain their evidence did not actually cover [9]. The Good Judgment Project’s internal findings echo this from the other direction: some of the accuracy gain attributed to “superforecasters” came not from superior domain knowledge but from disciplined question decomposition — breaking a vague, broad question into narrower, more resolvable sub-questions before assigning it a probability [10]. That is a procedural skill, transferable across domains, and it is a meaningfully different thing from being a better-informed expert in any one of them.
Institutional adoption of these methods has followed the incentives of the institutions doing the adopting, not necessarily the strength of the evidence behind each technique. Corporations tend toward scenario planning because it produces defensible strategic narratives without a scorable failure condition — a scenario team is rarely proven “wrong” in the way a calibrated forecaster can be. Intelligence and policy communities have shown more interest in tournament-style, track-record-based selection since the ACE program precisely because those communities are willing to accept scored accountability in exchange for the chance at a measurable edge. Both choices are rational responses to different institutional risk profiles, and the difference in adoption pattern is itself a data point about what each method is actually good for, separate from any claim about which one is more “scientific.”
None of these three programs was built to forecast technological discontinuity itself — a genuinely new capability arriving faster or later than any trend line implies — and that is the area where the historical record offers the least guidance and the most temptation to overreach. Scenario planning is structurally well suited to discontinuity, since it does not require assigning a probability to the break; a Delphi panel can be structurally biased toward the median expert’s timeline and therefore poorly suited to catching an outlier who turns out to be right early; and a tournament format depends on questions with fixed, near-term resolution dates, which by construction underrepresents slow-building structural change relative to fast, discrete geopolitical events.
Scenario (explicitly not a prediction): institutional forecasting over the next decade converges toward hybrid designs — scenario sets used to define the space of plausible futures, with tournament-style, track-record-selected forecasters or aggregation algorithms used to assign calibrated conditional probabilities within that space, echoing suggestions already present in RAND’s own current methodological literature listing Delphi, scenarios, and other techniques as complementary rather than competing tools [11].
Prediction, stated with a horizon and a disconfirmation condition: over the next five years, at least one major government or corporate forecasting unit will publicly report a formal track-record-based selection process for its in-house forecasters, modeled explicitly on the Good Judgment Project’s design, rather than relying solely on either open-panel Delphi elicitation or scenario workshops alone. Assumption: current interest in operationalizing superforecasting methods, visible in Good Judgment’s own commercial and institutional client work, continues at roughly its present pace [8]. Observable indicator: a published methodology document, procurement notice, or peer-reviewed evaluation naming calibration tracking as a formal selection criterion for forecasters. Disconfirmation condition: if by 2031 no major public-sector or Fortune 500 forecasting unit has published such a criterion, and instead credentialed-panel Delphi elicitation remains the dominant documented practice, the prediction is falsified.
Where the field goes from here is an open institutional question, not a settled one. What the history above supports is narrower and more durable: three specific programs solved three specific, non-interchangeable problems — how to elicit a defensible group estimate, how to prepare an organization for multiple futures without predicting one, and how to measure, for the first time, which forecasting practices actually track reality. Mistaking any one of the three for a general theory of the future is the error this history is meant to prevent.
Originally published at https://absolutedigitalpublishers.com/articles/from-origins-to-frontier-a-history-of-futures-methods-and-technological-scenarios.