Forecasting tournaments already measure how well humans and machines predict the future against each other. That record, not intuition, is the honest starting point for scenarios about 2035.

Forecasting tournaments turn prediction into an auditable practice: a question is written, a probability is logged, and an outcome later scores it. — Image prompt and art direction by Brecht Corbeel; generation pending.
Futures work is often presented as narrative art, but a quieter, measurable discipline has grown alongside it: forecasting tournaments that score predictions against outcomes and compute error with a scoring rule. This article separates what that record actually shows — superforecasters still beating frontier language models on a dynamic benchmark, prediction markets reaching tens of billions of dollars in volume while regulatory status stays contested, and a forecasting-research infrastructure that is still small and grant-dependent — from vendor claims, tournament organizers' framing, and the author's own analysis. It then builds three falsifiable scenarios for 2035, each with a stated horizon, explicit assumptions, observable indicators, and a disconfirmation condition, rather than a single narrative of AI-forecast ascendance or prediction-market mainstreaming.
Futures work has a reputation problem it partly deserves. Ask ten people what “futurism” means and most will describe a mood — chrome cities, flying cars, a keynote slide of exponential curves — rather than a method with a scoring rule. That reputation obscures a smaller, less glamorous practice that has existed for over a decade and produces something rare in this field: an auditable record of who was right, by how much, and when the world found out. Forecasting tournaments assign explicit numerical probabilities to well-defined future events, wait for the events to resolve, and score every forecaster with a proper scoring rule, usually the Brier score, where zero is perfect and higher is worse. That scored record — not the narrative art of scenario-writing — is the honest place to start an article about technological futures in 2035, because it is the only part of “futures methods” that can actually be checked against reality after the fact.
This article does three things. First, it lays out what the tournament record currently shows about forecasting itself: how good human superforecasters are relative to ordinary people and to frontier AI systems, how large and stable prediction markets have become as a pricing mechanism for future events, and how small and funding-dependent the research infrastructure behind all of this still is. Second, it is explicit about which claims in that record are measured facts, which are vendor or organizer framing, and which are this article’s own analysis. Third, it converts that record into three falsifiable scenarios for 2035 — not predictions of what will obviously happen, but conditional statements with a horizon, assumptions, an observable indicator, and a stated way each one could turn out to be wrong.
The modern tournament format traces to the Good Judgment Project (GJP), a research effort led by Philip Tetlock and Barbara Mellers at the University of Pennsylvania that competed in a four-year forecasting tournament run by the U.S. intelligence community’s research arm, IARPA, beginning in 2011 [9]. GJP recruited thousands of volunteer forecasters, had them assign probabilities to geopolitical questions, and used a simple statistical trick — identifying and regrouping the most accurate forecasters, later branded “superforecasters” — to consistently outperform the other tournament teams. Good Judgment’s own materials report that superforecasters were roughly 30 percent more accurate than a comparison group of intelligence analysts working with classified information, based on an independent study comparing GJP forecasts against the intelligence community’s internal prediction market [5]. That is a large claimed gap, and it comes from the tournament operator’s own retrospective account rather than an independent audit published outside Good Judgment’s materials — worth stating plainly as an organizational claim, not an adjudicated fact, even though it is broadly consistent with the peer-reviewed literature on GJP’s performance in the IARPA tournament [9].
The institutional afterlife of GJP is the Forecasting Research Institute (FRI), a nonprofit that Tetlock co-founded in late 2022 together with CEO Josh Rosenberg specifically to continue running large forecasting studies on consequential topics such as AI progress, biosecurity and geopolitics, and to translate forecasting methods into tools policymakers can actually use [4]. FRI now operates with a modest staff of roughly two dozen people plus academic affiliates and outside collaborators [3] — a research organization, not a government agency or a standing public institution, a distinction that matters for the scenarios later in this piece.
FRI’s most direct contribution to answering “can AI systems forecast as well as expert humans” is ForecastBench, a dynamic benchmark built specifically to defeat the usual complaint about AI evaluation — that a static test set eventually leaks into training data. ForecastBench instead generates and refreshes roughly a thousand forecasting questions on a rolling basis and tracks a public leaderboard, comparing predictions from human superforecasters, the general public, and large language models on a shared sample of 200 questions [1]. The headline result, as of the paper’s most recent revision, is unambiguous in direction if narrow in magnitude: superforecasters achieved a Brier score of 0.093 (lower is better), the general public scored 0.107, and the best-performing language model in the study scored 0.111 — a statistically significant gap favoring expert humans over the strongest model tested, with the gap between AI and ordinary members of the public much smaller than the gap between AI and trained superforecasters [1]. That is a measured result from a specific, dated study, not a permanent verdict on what AI can or cannot do; the paper’s own leaderboard is designed to be revisited as new models appear.
A second, independently run comparison points the same direction with more operational detail attached. Metaculus, a forecasting platform that runs recurring tournaments pitting bot-authored forecasts against a panel of vetted human “Pro” forecasters, has published head-to-head scores for its AI forecasting tournament by quarter. The best-performing bot team scored -11.3 in Q3 2024 and -8.6 in Q4 2024 in head-to-head comparison against the Pro forecaster panel, where a score of zero would indicate equal accuracy and negative scores indicate the bots underperforming; the Pro human panel had, as of that data, beaten the bot panel in every quarterly comparison run to date, though the gap narrowed from one quarter to the next [2]. Two things should be separated here. The fact is: in the specific tournaments measured through late 2024, professional human forecasters outscored the best available bot teams, and the score gap was shrinking. The analysis — and this is this article’s own reading, not a claim made by either Metaculus or FRI — is that a shrinking gap over only two or three quarters is too short a trend to extrapolate confidently in either direction; model capability, prompting technique, and the specific question mix each quarter all move the number, and none of the public sources isolate how much of the narrowing is due to genuinely better forecasting reasoning versus better tournament-specific engineering by bot-team operators.
It is also worth being explicit about what neither source claims. Neither ForecastBench nor the Metaculus tournament data supports a claim that AI forecasting has already matched or surpassed human experts as of 2026; both show a real but incomplete convergence, measured on specific question sets that may not represent all forecasting tasks a policymaker or a company would actually care about.
A related but distinct thread in modern futures method is the prediction market — a venue where people trade contracts whose payout depends on a future event, so that the market price is interpretable as a crowd-aggregated probability. Unlike tournament forecasting, which is a research exercise, prediction markets are now a functioning commercial industry with real trading volume. Kalshi and Polymarket together accounted for roughly 97.5 percent of prediction-market activity in 2025, with Kalshi’s notional volume reported at approximately $23.8 billion — more than 1,100 percent year-over-year growth — and Polymarket at roughly $22 billion, while all other platforms combined handled about $1.25 billion [6]. That volume figure is a market-data claim from a trade-press source, not an audited public filing, and it should be read with that caveat; it is nonetheless a large enough number, and consistent enough across independent reporting, to treat as a real order of magnitude rather than a rounding error.
Regulatory status is where “adoption” and “legitimacy” pull apart, and this is exactly the kind of place where a forecasting article should resist collapsing disagreement into a single verdict. The Commodity Futures Trading Commission (CFTC) holds exclusive federal jurisdiction over these contracts, treating them as swaps traded on CFTC-designated exchanges, and has issued an advance notice of proposed rulemaking signaling that formal federal rules are coming [7]. At the same time, several states have pushed back hard: Arizona filed a twenty-count criminal case against a Kalshi-affiliated entity, Illinois’s gaming board sent cease-and-desist letters to exchange operators, and Connecticut asserted that sports-related event contracts amount to illegal gambling under state law [7]. A federal appeals court sided with federal preemption in one specific case in 2026, but the underlying conflict between federal exchange law and state gambling law remains open, and further litigation, including a possible path to the Supreme Court, is plausible [7]. Separately, the CFTC has been actively encouraging exchanges to formalize partnerships with sports leagues, which reads as an agency trying to normalize a category it is still writing rules for rather than one it has already fully legitimized [8].
Put plainly: prediction markets have institutional-scale trading volume and CFTC-regulated exchanges already operating with self-certified contracts numbering in the thousands, covering elections, sports, macroeconomic indicators and weather [7]. What they do not yet have is settled legal status, an agreed federal-state division of authority, or documented use by a government body as a formal policy-decision input at any significant scale. Those are two different claims, and vendor and platform commentary sometimes elides them; this article is keeping them separate on purpose.
Before turning to scenarios, it is worth naming the categories this article has been using, because the futures-studies literature is unusually prone to blurring them. A fact here means a number or event this article personally verified against a live source: the ForecastBench Brier scores, the Metaculus quarterly head-to-head scores, the reported 2025 trading volumes, the CFTC’s rulemaking notice. A vendor or organizer claim is a number or characterization that comes from an interested party describing its own product or tournament, such as Good Judgment’s 30-percent accuracy claim relative to intelligence analysts, or a platform’s self-reported market share — plausible, sourced, but not independently adjudicated. Analysis is this article’s own interpretation connecting facts, such as the observation that a two-quarter narrowing trend in bot-versus-Pro scores is too short to extrapolate. Scenario and prediction are explicitly conditional statements about 2035, laid out below, and are never presented as facts about what will happen.
Each scenario states a horizon, the assumptions it depends on, what would count as an observable indicator along the way, and — critically — an explicit condition under which the scenario should be considered wrong, not just delayed.
Horizon: by year-end 2035, evaluated against whichever benchmark descended from ForecastBench (or a successor with comparable methodology) is still actively maintained at that date.
Assumptions: that a comparable dynamic, leak-resistant forecasting benchmark continues to be run and published at least annually; that “parity” is defined as the best AI system’s Brier score falling within the published year-to-year variance band of the superforecaster cohort’s own score, not merely closing most of the numerical gap once.
Observable indicators between now and then: successive ForecastBench-family leaderboard updates showing the AI-to-superforecaster gap narrowing in at least three of four consecutive reporting periods; Metaculus or comparable tournament head-to-head scores crossing from negative to at or near zero and staying there across multiple quarters rather than a single quarter; independent replication of any parity claim by a second benchmark operator, since a single organization’s leaderboard is a thinner form of evidence than two disagreeing groups converging.
Disconfirmation condition: if by 2035 the best documented AI system’s score on the maintained benchmark remains outside the superforecaster cohort’s variance band, or if the gap has stopped narrowing for multiple consecutive reporting periods (a plateau rather than continued convergence), this scenario is falsified — not “still developing,” but wrong as stated.
Horizon: by year-end 2035.
Assumptions: current CFTC rulemaking eventually resolves into a stable federal framework (rather than remaining in open litigation indefinitely); at least one prediction-market operator remains solvent and CFTC-registered through that period; “mainstream institutional adoption for policy” is defined narrowly as a government body or its published advisory process explicitly citing a specific prediction-market price as a named input to a specific decision, in a public document — not merely that officials or advisors are reported to watch markets informally.
Observable indicators: continued growth in CFTC-self-certified contract categories beyond elections, sports and macro indicators reported in current disclosures [7]; a published federal rule (not just an advance notice) settling the state-preemption question; any public agency methodology document, budget justification, or regulatory impact analysis that names a specific market price as an evidentiary input.
Disconfirmation condition: if by 2035 no public government document can be found naming a specific prediction-market price as a decision input — even if trading volume keeps growing and even if officials continue informally referencing markets in interviews — this scenario is falsified. Volume and legal survival are not the same as documented institutional use; this article treats “informally watched by some officials” as insufficient evidence for “mainstream institutional adoption.”
Horizon: by year-end 2035, assessed by counting actively funded forecasting-research organizations comparable in scope to FRI (multi-year staff, published methodology, public benchmark or tournament output) and their combined disclosed annual funding.
Assumptions: that organizations in this category continue to disclose staff size and funding sources at roughly the level FRI currently does [3]; that “expansion as a field” means growth in the number of independently operating organizations of this kind, not just growth in the budget of any single one.
Observable indicators: new forecasting-research nonprofits or university centers publishing their own dynamic benchmarks or tournament series, comparable in rigor to ForecastBench, appearing over the coming decade; continued or growing philanthropic and government grant funding directed at this research category rather than a single major funder’s support lapsing without replacement; government agencies (following the pattern of IARPA’s original tournament) commissioning new forecasting tournaments of their own.
Disconfirmation condition: if by 2035 the number of comparably rigorous, independently operating forecasting-research organizations has not grown beyond roughly today’s handful, or if today’s organizations have shrunk, merged, or gone dormant for lack of funding, this scenario is falsified — the field will have stayed a niche academic-adjacent activity rather than becoming a durable, expanding research infrastructure.
It is worth pausing on a limitation of this entire evidentiary base before moving to scenarios, because it bears directly on how much weight the 2035 forecasts below can carry. Tournament forecasting and prediction markets both work best on questions with a clear, checkable resolution: did a named event happen by a named date, yes or no, or within a stated numeric range. That is precisely why IARPA’s original tournament, ForecastBench, and Metaculus’s quarterly comparisons all lean on geopolitical events, economic indicators, and dated technology milestones — categories where “did it happen” is rarely ambiguous. Long-range, structural questions about technology and society, the kind futurists are usually asked to answer, resist this format much more. “Will AI forecasting matter to how governments plan” is not a single resolvable event; it has to be decomposed into narrower, checkable sub-questions before a tournament or a market can price it at all, which is exactly what this article’s three scenarios attempt to do. That decomposition is itself a methodological choice, and a different analyst decomposing the same broad question could reasonably choose different sub-questions and arrive at different indicators. Readers should treat the scenarios below as one defensible decomposition, not the only possible one.
This limitation also explains why the article separates “tournament-measured fact” from “scenario” so insistently. A Brier score is only as meaningful as the resolution criteria behind the question it is scoring, and a benchmark built from a thousand well-defined questions says something real about probabilistic calibration on that kind of question — it does not automatically transfer to open-ended strategic judgment, let alone to questions no one has yet learned how to phrase in resolvable form. Vendors selling AI forecasting tools sometimes elide this distinction, presenting a benchmark score as evidence the tool can handle any forward-looking business or policy question a client might bring; the tournament literature itself does not support that generalization, and this article is not making it either.
A conventional futures piece would fold these three threads into one story — “AI will out-forecast humans and price the future through liquid prediction markets used by governments.” That story is more satisfying to read and less true to what the record supports. The measured facts above show partial, quarter-by-quarter convergence in one domain (AI-versus-human forecasting accuracy), large but legally contested growth in another (prediction-market volume), and a small, still-fragile research base underneath both. Any one of these three could plausibly reverse without the others moving at all: a legal setback could freeze prediction-market growth while AI forecasting accuracy keeps improving; funding for tournament research could dry up even as commercial prediction markets keep scaling; AI systems could plateau below superforecaster accuracy indefinitely even as markets and tournament infrastructure both keep growing. Treating them as three separable, falsifiable claims — each with its own horizon and its own way of being proven wrong — is a more honest use of “futures methods” than a single unified prophecy, and it is closer to what the tournament tradition this article opened with actually trained its participants to do: assign a number, write down what would change your mind, and wait to be scored.
Two practical implications follow directly from the record above, without requiring any 2035 scenario to resolve first. First, an organization deciding whether to rely on AI-generated forecasts today, rather than in 2035, is choosing a tool that the best current public benchmarks still place behind trained human superforecasters, though not far behind ordinary informed judgment [1]. Treating an AI forecast as a categorical replacement for expert human judgment is not supported by the record as it currently stands; treating it as a supplement that approaches lay-person-level accuracy at far lower cost and higher speed is closer to what the data shows. Second, an organization considering prediction-market prices as a policy input today is adopting a mechanism with real trading depth but genuinely unsettled legal footing [7] [8] — a reasonable experimental input, not yet a mechanism with the kind of institutional precedent that would make its use unremarkable. Both of those are analysis, not fact, and both are exactly the kind of claim a forecasting tournament would ask its participants to state a probability on, rather than assert as settled.

Figure 1. A calibration wheel checks whether events assigned a given probability happen about that often — the arithmetic test that separates a forecaster's confidence from their accuracy. — Image prompt and art direction by Brecht Corbeel; generation pending.

Figure 2. Head-to-head tournament scores are the only place this comparison is actually measured: through late 2024 the strongest bot teams still trailed professional human forecasters, though the gap narrowed each quarter. — Image prompt and art direction by Brecht Corbeel; generation pending.

Figure 3. Prediction markets price event probabilities the way an exchange prices anything else; whether that pricing becomes a policy input or stays a retail curiosity is still contested ground, not a settled fact. — Image prompt and art direction by Brecht Corbeel; generation pending.

Figure 4. The tournaments that make this kind of scoring possible are a small, grant-dependent research infrastructure, not a mature public institution — whether that infrastructure grows or contracts is itself an open question. — Image prompt and art direction by Brecht Corbeel; generation pending.

Figure 5. Every forecast in this article carries a stated horizon and a disconfirmation condition; a prediction without a date it could fail by is not a forecast, it is a mood. — Image prompt and art direction by Brecht Corbeel; generation pending.
Originally published at https://absolutedigitalpublishers.com/articles/futures-methods-and-technological-scenarios-in-2035-scenarios-signals-and-falsifiable-predictions.