Four boards, not one forecast

Most writing about where the frontier-model contest ends up by 2035 is a single-track story dressed as analysis. One version has three or four labs racing forever, leaderboards deciding bragging rights, and today’s leaders staying on top because they got there first. Another has the field narrowing to whichever two or three organisations can still afford a frontier training run, with comparison moving on from leaderboard position to which vendor’s models stay deployed. Both are coherent. Neither is more than a preference dressed as a forecast, because the evidence available in August 2026 does not discriminate between them.

The honest alternative is scenario analysis, with a specific discipline: name the smallest number of drivers that are both consequential and genuinely uncertain, cross them, state the mechanism behind each resulting cell, and commit in advance to the observations that would identify which cell the industry is in and the observations that would rule a cell out. Done properly it produces no favourite — that is the intended output, not a failure to reach one.

This article uses two axes. Axis A asks how many organisations remain within a genuinely competitive distance of the frontier: a plural field of labs, or consolidation to two or three dominant organisations. Axis B asks what the public treats as the primary signal of who is actually ahead: composite benchmark and leaderboard position, which has organised most public comparison to date, or measured real-deployment outcomes — production task completion, enterprise retention, revenue capture — that some of this field’s newer instruments are already built to supply. Two frequently cited related questions, whether open-weight models are closing the capability gap with closed frontier models and whether providers are converging on similar capability or continuing to differentiate, are treated below as consequences of these two axes rather than independent drivers, a reduction argued for explicitly rather than assumed.

ADVERTISEMENT

The documented present

Fact. The population of organisations near the top of the field has been reshuffling rapidly, not settling. Menlo Ventures’ enterprise survey, published 9 December 2025, put Anthropic at roughly forty percent of enterprise large-language-model API spending, more than triple its twelve percent share in 2023; OpenAI’s share fell to twenty-seven percent from fifty percent; Google climbed to twenty-one percent from seven percent; the three together hold eighty-eight percent of enterprise usage, the remainder spread across Meta’s Llama, Cohere, Mistral, and smaller providers [8]. An independent measure agrees on direction from a different source: Ramp’s payments-based index, reported in May 2026, recorded Anthropic passing OpenAI in business-to-business adoption for the first time, thirty-four-point-four percent to thirty-two-point-three, with Anthropic’s share reported as quadrupling over the preceding year against OpenAI’s three-tenths-of-a-point gain [12]. A revenue survey and a transaction index agreeing on direction, from two different methodologies, is reasonable confidence to place in one finding: market position among the current leaders is genuinely contested, not fixed.

Fact. The capital cost of training at the frontier is rising fast enough to plausibly narrow the field on its own. Epoch AI’s cost model estimates the amortised hardware and energy cost of the largest training runs has grown roughly two-point-four times per year since 2016, with a cloud-rental-based method giving a comparable two-point-six times per year, and projects the largest runs will cost more than a billion dollars by 2027, restricting frontier-scale training, on the authors’ own reading, “to only a few large organizations” [5]. Epoch’s trend dashboard separately reports frontier training compute growing roughly five times per year since 2020 — doubling near every five months — partly offset by pretraining compute efficiency improving roughly three times per year, so net capital requirement, while still rising steeply, rises somewhat more slowly than the raw compute curve alone suggests [6].

Write Ctrain(t)C_{\text{train}}(t) for the cost of the largest training run in year tt. Epoch’s estimate is well approximated by simple exponential growth,

Ctrain(t)C0gtt0, C_{\text{train}}(t) \approx C_0\, g^{\,t-t_0},

with gg between roughly two-point-four and three-point-five depending on method and window [5, 6]. The form matters more than the exact rate: exponential cost growth does not by itself predict consolidation. It predicts a shrinking count of organisations able to clear an ever-rising bar only if the pool of adequately capitalised entrants grows more slowly than gtg^{t}. Whether that pool keeps pace is a question about capital markets and state investment, not something the cost curve alone answers — which is exactly why axis A stays open below rather than settled by this one fact.

Fact. The infrastructure layer beneath the frontier labs is already structured in ways a regulator has flagged as a concentration risk independent of who wins on capability. The U.S. Federal Trade Commission’s Office of Technology completed a Section 6(b) study of three cloud-service-provider-to-AI-developer partnerships — Microsoft and OpenAI, Amazon and Anthropic, and Google and Anthropic — publishing its staff report on 17 January 2025. The report documents the equity and revenue-sharing rights the cloud providers hold, the consultation, control and exclusivity provisions attached to their investments, and flags that the partnerships could affect access to computing resources and engineering talent, raise switching costs for the AI-developer partners, and give the cloud-service partners access to sensitive technical and business information not available to others [7]. This is a regulatory finding about deal structure, not about which model is better, but it bears directly on axis A: if compute access and switching costs are structurally tied to a small number of cloud relationships, the number of organisations that can plausibly stay at the frontier is bounded by something other than research talent alone.

ADVERTISEMENT
A brushed-steel market-positioning wall with a tight cluster of manila index cards, one card caught rising clear of its magnet with a bright sliver of steel showing where it sat, corner still curled
Figure 1. A shrinking field looks like this before it looks like anything else: one position lifted off the board, the gap still bright and empty behind it.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Fact. Public benchmark and leaderboard comparison remains the field’s dominant reference point, and it has a documented reliability problem. Chatbot Arena, introduced by Chiang and colleagues, collects crowdsourced pairwise human-preference votes between anonymised models and estimates rankings with Bradley-Terry-style statistical methods; the platform accumulated more than two hundred forty thousand votes and became one of the most-referenced LLM leaderboards among developers [1]. A later audit, Singh and colleagues’ “The Leaderboard Illusion,” examined roughly two million battles across two hundred forty-three models and forty-two providers and found systematic distortions: providers can test private model variants before public release and selectively disclose only favourable results — the paper documents twenty-seven private variants Meta tested ahead of the Llama 4 release — and sampling is asymmetric, with two large closed-model providers alone receiving roughly nineteen to twenty percent of arena battles each while eighty-three open-weight models combined received under thirty percent, an imbalance the authors show can inflate apparent performance gains well beyond what the same additional data would buy a fairly sampled model [2]. Whether this is a temporary wrinkle or evidence that leaderboard-style comparison cannot scale to a crowded, high-stakes field is precisely what axis B is built to track.

Fact. A rival signal, built around real deployment outcomes rather than composite scores, is under construction. OpenAI’s GDPval, described by Patwardhan and colleagues, evaluates models against one thousand three hundred twenty tasks drawn from the actual work of forty-four occupations across the nine sectors contributing most to U.S. GDP, built from the representative work of professionals with roughly fourteen years of average experience and graded in blind comparison against those professionals’ own deliverables; the authors report frontier performance improving roughly linearly and approaching expert quality, with models completing tasks on the order of a hundred times faster and cheaper than the human experts sampled — a vendor-reported claim, not an independently replicated one [3]. The boundary between “benchmark” and “deployment outcome” is already blurring: the Artificial Analysis Intelligence Index, an independently operated composite of nine evaluations across agentic, coding, general-knowledge and scientific-reasoning categories, assigns GDPval’s own automated variant roughly a third of the weight in its “agentic” category, folding a deployment-shaped evaluation into what is still reported as one leaderboard number [11].

A mechanical adding machine on a research bench with a curl of paper tape spilling from its print head, the tape still feeding out and not yet torn off, a ledger book open beside it
Figure 2. What a rising cost curve looks like up close: the tape keeps feeding, and no one has torn it off yet to read the total.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Fact. Independent trend evidence shows capability rising on measures that are neither pure leaderboard score nor pure deployment outcome. METR’s time-horizon study, using a fifty-percent task-completion-probability metric on real software tasks, reports frontier time horizon roughly doubling every seven months since 2019, with tested frontier models reaching roughly fifty-minute horizons at publication and the authors flagging uncertain generalisation beyond the studied distribution [9]. Separately, Guo and colleagues’ peer-reviewed account of DeepSeek-R1 showed a reasoning policy trained by reinforcement learning against verifiable rewards, without human-annotated traces, raising AIME 2024 pass-at-one accuracy from 15.6 percent to 77.9 percent over training — a result from an open-weight lab, published in Nature, that both demonstrates a general mechanism (deliberation is a trainable policy) and shows how quickly an organisation outside the traditional three-way U.S. contest can reach frontier-tier performance on one measure [10].

Fact. Stanford’s 2026 AI Index Report, the seventh edition of an independently steered annual survey, documents both convergence and its opposite inside the same period. It reports that “U.S. and Chinese models have traded the lead multiple times since early 2025,” that DeepSeek-R1 briefly matched top U.S. models in February 2025, and that by March 2026 the leading model held only a two-point-seven-percent advantage over the assessed field — convergence, on a composite score. The same report finds hallucination rates across twenty-six top models ranging from twenty-two to ninety-four percent, GPT-4o’s accuracy on a knowledge-versus-belief test falling from 98.2 to 64.4 percent, and DeepSeek R1’s falling from over ninety percent to 14.4 percent — divergence, on a different measure, same period [4]. Both findings are accurately reported; they answer different questions, this article’s central methodological point, stated here in miniature before the two axes are built from it.

Why these two axes, and why the open-weight gap and capability convergence are not a third and fourth

An axis earns inclusion by being both consequential and genuinely uncertain. Axis A is consequential because its two ends imply different economics, different regulatory exposure, and different failure modes: a plural field means buyers retain real alternatives and no single outage or safety incident removes most of the market’s capacity at once, while a narrow field concentrates both benefit and risk in a small number of balance sheets and cloud relationships. It is uncertain because the evidence points both ways inside the same eighteen months: a cost curve growing two-and-a-half to three-and-a-half times a year and a regulator’s finding of structurally advantaged cloud partnerships [5, 7] point toward narrowing, while three-way market-share churn large enough for the leading vendor to change twice in three years [8, 12] and a peer-reviewed frontier-tier result from outside the traditional three [10] point toward continued plurality.

Axis B is consequential because it decides what a claim of leadership actually means and to whom it is addressed — a leaderboard number chiefly addresses developers, procurement teams and the press; a deployment-outcome number addresses buyers who have already integrated a model and are deciding whether to keep it. It is uncertain because the evidence points both ways too: composite benchmark and arena-style comparison remains the field’s most-cited reference point [1], and an independent audit of that same reference point has documented private pre-release testing and lopsided sampling serious enough to move apparent rankings by more than the underlying capability difference [2], while a deployment-outcome benchmark built from real occupational tasks is maturing quickly enough to already be folded into a composite index as one input among several [3, 11].

ADVERTISEMENT

A convenient way to see what axis B actually claims is to write the signal a given comparison reports as a blend:

S(t)=w(t)B(t)+(1w(t))D(t), S(t) = w(t)\, B(t) + \bigl(1 - w(t)\bigr)\, D(t),

where B(t)B(t) is a provider’s position on a composite benchmark or leaderboard index at time tt, D(t)D(t) is its position on a deployment-outcome measure such as production task completion or retained enterprise share, and w(t)[0,1]w(t) \in [0,1] is the weight public comparison, procurement practice and press coverage place on the benchmark term. Axis B is precisely a claim about where w(t)w(t) is heading: toward one, where the scenario names below carry the word “Board,” or toward zero, where they carry the word “Ledger” instead. No source cited in this article reports w(t)w(t) directly as a number, which is the honest reason this is an axis and not an already-settled fact.

Two frequently asked questions about 2035 are not independent of these two axes, and folding them in as a third and fourth would double-count evidence already used above.

Whether open-weight models are “closing the gap” with closed frontier models is really axis B asked about one class of providers rather than all of them, and it returns different answers depending on which signal does the measuring. On the signal that behaves like axis B’s benchmark end, open-weight entrants have moved fast: Guo and colleagues’ DeepSeek-R1 result reached frontier-tier reasoning performance on a public math benchmark from an open-weight lab [10], and Stanford’s Index reports that same model briefly matching the top U.S. models in February 2025, with the field’s leading model retaining only a narrow lead over the assessed pool thirteen months later [4]. On the signal that behaves like axis B’s deployment-outcome end, the same period tells a different story: Menlo’s enterprise survey found open-source models’ share of enterprise LLM API spending fell from nineteen percent to eleven percent even as the benchmark story above was unfolding [8], and the Leaderboard Illusion audit’s sampling-asymmetry finding — eighty-three open-weight models sharing under thirty percent of measurement battles, against roughly two-fifths for two closed providers alone — means even the benchmark-side “gap” numbers already in circulation were measured on an uneven instrument [2]. “Is the open-weight gap closing” is not a fifth question with its own independent evidence base. It is axis B, and it will keep returning different answers depending on which signal a given source used to ask it — which is the finding, not an evasion of one.

An open oak card-catalog drawer full of printed index cards, a brass divider tab caught lifted halfway out of its slot at the point where two sections meet, cards from each side leaning together across the gap
Figure 3. The line between one kind of model and another is a tab that can be lifted; caught mid-lift, the two sections are already leaning into each other.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Whether providers “converge” toward similar capability or continue to differentiate is likewise mostly a readout of how many organisations stay inside the frontier band (axis A) crossed with which signal judges convergence (axis B). The same Stanford Index that reports a two-point-seven-percent gap between the leading model and the field on one composite measure also reports a hallucination-rate spread from twenty-two to ninety-four percent across twenty-six models, and directly opposite year-over-year movement between two named models on a knowledge-versus-belief test, in the same report [4]. A single publication saying “converged” and the same publication saying “diverged” can both be true of the same providers in the same period; which claim a reader encounters depends entirely on which signal, and which competitive set, produced it. Treating convergence as its own axis would mean re-deriving axis B’s question under a different name rather than adding real information.

Four scenarios toward 2035

Crossing the two axes gives four cells. None is named as the likely outcome; each is named for the instrument in this article’s own room that best captures its logic — “Board” for scenarios where composite index or leaderboard position stays the dominant public signal, “Ledger” for scenarios where measured deployment outcome takes over that role.

Scenario one: Open Board — plural field, benchmark-led signal

Mechanism. Enough organisations keep clearing an ever-rising capital bar — helped by pretraining efficiency gains running near three times a year, partly offsetting a raw compute-cost curve growing two-and-a-half to three-and-a-half times a year [5, 6], and helped by entrants from outside the traditional three-way U.S. contest reaching frontier-tier results on at least some measures [10] — that a genuinely plural field survives. At the same time, public comparison keeps organising around composite index and leaderboard position, because that reference point is already deeply embedded in press coverage, procurement habits and system-card conventions, and because no rival signal has yet displaced it despite documented reliability problems [1, 2].

Horizon. Recognisable by 2029–2030; a stable configuration plausible through 2035.

Assumptions. Efficiency gains keep pace with, or outrun, capital-cost growth for enough entrants to stay solvent; no regulatory action forces structural separation of the cloud-AI partnerships the FTC has already flagged [7]; index operators remain the reference point despite documented gaming.

Observable indicators. First, five or more organisations remain within a modest gap of the top composite-index score through 2030. Second, system-card coverage and vendor announcements continue to lead with index or leaderboard position as the headline comparison. Third, at least one organisation from outside the traditional three-way U.S. contest stays inside the top tier of a major public index across multiple index revisions.

Disconfirmation. Falsified if, by 2030, the count of organisations within a modest gap of the top index score falls to two or fewer for two consecutive years, or if major system-card coverage has visibly shifted to leading with deployment-outcome figures instead.

Scenario two: Open Ledger — plural field, outcome-led signal

Mechanism. The supply side stays plural for the same reasons as scenario one, but buyers stop deciding between vendors chiefly by index score and act instead on measured production outcomes, rewarded because the benchmark signal has a documented gaming problem [2] and deployment-outcome measurement is maturing into a usable alternative [3, 11]. With several labs still viable, which one wins on outcome measures keeps rotating, consistent with the churn already visible between three providers in eighteen months [8, 12].

Horizon. Recognisable shift by 2029–2030; steady state by 2035.

Assumptions. Deployment-outcome measurement — task-completion suites, spend-tracking indices — becomes standardised and trusted enough to be cited routinely rather than as one-off studies; enough labs remain solvent to keep the ranking genuinely contested rather than settling on one incumbent.

Observable indicators. First, vendor announcements and independent coverage lead with production task-completion or enterprise-retention figures rather than composite index scores. Second, spend-tracking indices in the style of the Menlo and Ramp measures cited above recur on a regular publication cadence rather than appearing once. Third, no single provider holds a majority of a tracked market for three consecutive years.

Disconfirmation. Falsified if composite index position remains the number vendors and press lead with through 2030, or if deployment-outcome indices fail to recur past their first one or two editions.

A small brass beam balance on a research bench, one pan holding a stack of printed benchmark scorecards and the other holding a handful of poker-chip-style outcome tokens, the beam caught tilted and still swinging
Figure 4. The instrument for weighing one kind of evidence against another; the beam is still swinging, not resting on either pan.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Scenario three: Narrow Board — consolidated field, benchmark-led signal

Mechanism. The capital-intensity curve and the structural advantages the FTC documented in cloud-AI partnerships [5, 7] squeeze frontier-scale training down to a small, stable set of organisations. Because that set is small and each member can still afford to optimise for public index position relative to its overall budget, and because press and procurement habits built around leaderboard comparison are sticky, composite benchmark position stays the dominant public signal even though only a handful of organisations are left to contest it.

Horizon. Recognisable by 2029–2031; stable by 2035.

Assumptions. No efficiency breakthrough beyond Epoch’s roughly three-times-a-year trend reopens frontier-scale training to a wider field [6]; cloud-AI partnership structures, or comparable successors, continue conferring compute-access advantages rather than being unwound.

Observable indicators. First, the count of organisations publishing a frontier-scale system card near the top of a major index falls to three or fewer and stays there for three consecutive years. Second, at least one reported single-training-run cost crosses a billion dollars, as Epoch’s trend projects for 2027 [5]. Third, cloud-AI partnership terms — equity stakes, exclusivity, revenue share — remain in place or deepen rather than being unwound by regulatory action.

Disconfirmation. Falsified if the count of organisations near the top of a major index rises back above three for three consecutive years, or if a materially larger efficiency jump than Epoch’s software-efficiency trend reopens frontier-scale training to a wider field.

Scenario four: Narrow Ledger — consolidated field, outcome-led signal

Mechanism. The same capital and partnership concentration narrows the field as in scenario three, but the contest among survivors stops being about index rank and becomes about which incumbent keeps its deployed footprint. Menlo’s data already shows one provider’s coding-tool share running far ahead of its nearest rival inside a fast-rotating three-way market [8], and the FTC report flags switching costs as a structural feature of the cloud-AI partnerships underneath the concentration [7]. With only two or three plausible vendors, procurement is decided on integration and reliability, not incremental index points, because the buyer’s realistic alternative set is small enough to evaluate directly.

Horizon. Recognisable by 2030–2031; a stable configuration plausible by 2035.

Assumptions. The capital and partnership concentration of scenario three holds; buyers standardise on outcome- or switching-cost-based procurement once the vendor set is small enough to compare directly.

Observable indicators. First, the same narrow-field indicators as scenario three. Second, enterprise-adoption and spend-tracking reporting becomes the primary cited comparison in trade coverage, displacing index scores. Third, switching-cost evidence — contract lock-in periods, exclusivity terms — becomes a standard element of market analysis rather than a one-off regulatory finding.

Disconfirmation. Falsified if, despite a narrow field, composite index position remains the primary cited comparison through 2031, or if the field widens back out under either plural-field scenario’s own disconfirmation terms.

What all four share, and the possibility neither axis names

Three things hold across every cell, and are the safest things to build practice on regardless of which one obtains. First, some form of public comparison survives in all four — even Narrow Ledger, the most consolidated and outcome-oriented cell, still implies someone comparing the two or three survivors, because buyers with real alternatives, however few, need a basis for choosing between them. Second, the tension documented above between benchmark-side and deployment-side evidence on the open-weight question does not resolve in any cell; it is a structural feature of having two signals rather than one. Third, none of the four requires a capability plateau or discontinuity — each is compatible with capability continuing to rise at something like the rates METR and Epoch already document [9, 6], because the axes describe market structure and measurement practice, not the rate of technical progress.

The far end of the capability-tracking bench's brass rail, its last graduation mark catching the light, and a single turned-brass token resting on bare walnut just beyond the rail's end, off the graduated length entirely
Figure 5. Nothing here rules out a fifth thing neither axis names; a token resting past the last mark is what that looks like before anyone has decided it belongs on the rail at all.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

The four scenarios share a blind spot too, worth naming rather than hiding. All four assume the two axes here remain the two that matter. A large enough exogenous shock — a change in export or compute-access policy that redraws which organisations can buy frontier-scale hardware at all, a safety or security incident serious enough to trigger emergency restrictions outside the gradual processes every scenario above assumes, or a shift in what “a model” even is, toward a persistent multi-agent system evaluated by neither a single index score nor a single deployment-outcome figure — would not fit cleanly into any of the four cells, because it would move both axes at once, abruptly, rather than along the gradual paths each scenario traces.

Three predictions, stated separately from the scenarios

Prediction one. Horizon: end of 2029. At least one deployment-outcome-style evaluation, in the style of GDPval or a comparable task-completion suite, will be cited as a standard, recurring reference point by two or more major providers in their own system-card or release communications, not only by independent index operators. Assumption: no single catastrophic measurement failure discredits deployment-outcome evaluation in its first few cycles. Indicator: system cards or release posts from two or more providers citing a task-completion score on a recurring basis. Disconfirmed if by the end of 2029 provider communications still cite only composite benchmark or arena-style scores as their headline claim.

Prediction two. Horizon: end of 2030. The count of organisations publishing a system card within a modest gap of the top score on a major composite index will not fall below four, because efficiency gains and non-U.S. entrants are already offsetting capital-cost growth faster than the raw cost curve alone suggests [6, 10]. Assumption: no regulatory action or capital-markets shock removes two or more current top-tier organisations from frontier-scale training. Indicator: published index rosters and their score gaps at each major revision. Disconfirmed if the top-tier count falls to three or fewer and stays there through 2030.

Prediction three. Horizon: end of 2031. Regardless of which scenario the industry tracks toward, at least one further independent audit in the style of the Leaderboard Illusion paper will document a comparable measurement distortion in whichever signal, benchmark or deployment-outcome, has become dominant, because neither signal type has yet been shown resistant to the incentive that produced the first such audit [2]. Assumption: independent auditors retain access to the data needed. Indicator: a published audit, from a group without a commercial stake in the audited signal, documenting a specific distortion mechanism. Disconfirmed if no such audit of the dominant signal appears by the end of 2031, whether because the signal proved robust or access to audit it was foreclosed.

What to take away

The refusal to name a favourite among these four is the substantive claim, not a hedge around one. In August 2026 the evidence is genuinely split on both axes at once: a capital-cost curve growing two-and-a-half to three-and-a-half times a year sitting beside three-way market-share churn large enough to change the leading vendor twice in eighteen months; a leaderboard tradition that remains the field’s most-cited reference point sitting beside a published audit of that same tradition’s sampling and disclosure practices; a composite index reporting a two-point-seven-percent gap between the leader and the field in the same report that documents a hallucination-rate spread running from twenty-two to ninety-four percent. Anyone reporting a confident single future for the frontier-model landscape in 2035 is reporting which of these four they would bet on, not what the current record shows. The more useful and less satisfying discipline is the one this article tried to practise throughout: know which signal to watch, and have said in advance, on the record, what each one finding would mean.