Two axes, not four separate bets

Ask five people working on multimodal AI where the field will be in 2035 and the answers cluster around four separate bets, argued as though independent. Will training converge on a single model that learns text, image, audio, video and action together from the start, the way GPT-4o’s own account describes training “a single new model end-to-end across text, vision, and audio, meaning that all inputs and outputs are processed by the same neural network” [1] — or will the modular pattern, a frozen backbone plus swappable per-modality adapters, kept modular because joint-from-scratch multimodal training is prohibitively costly, keep being the pragmatic default [6]? Will real-time video and audio understanding get reliable enough that a robot arm can act on it in an ordinary room, the way Gemini Robotics and π₀ already act on it inside their own evaluation settings [7, 8] — or does a benchmark built specifically to test generalization, which found that “no model demonstrates consistent generality” across unseen domains, describe a durable ceiling rather than a growing pain [9]? Will video and 3D generation stop being something a viewer forgives and start being something a studio ships without a human pass, when the best current video model still clears joint semantic-and-physics adherence on fewer than one in four of the hardest test cases [11]? And will one any-to-any model replace the specialist pipeline, or does asking one system to cover six capability regimes at once turn out to be exactly the setup that produces “catastrophic knowledge degradation under domain transfer” [9]?

Argued one at a time, these read as four independent coin flips. They are not. The first and the last are the same underlying question approached from two directions — training method and deployment interface — and the second and third are also the same underlying question approached from two directions — perception and generation. This article uses two axes instead of four. Axis A, architecture, asks whether the field converges on natively trained, single-model any-to-any systems or stays modular. Axis B, reliability, asks whether multimodal perception and generation both cross into dependable open-world operation or stay confined to curated conditions. Crossing them gives four scenarios, each with a horizon, assumptions, observable indicators and an explicit disconfirmation condition. None is named as the likely outcome.

The documented present

Fact. OpenAI’s own account of GPT-4o states plainly that the company “trained a single new model end-to-end across text, vision, and audio,” with all inputs and outputs processed by that one network — a genuine architectural claim, since it is offered in contrast to a prior pipeline of separate transcription, language and image models chained together [1]. The same account reports response latency low enough to sustain natural spoken conversation.

ADVERTISEMENT

Fact. Google’s Gemini family was built the same way: designed as a multimodal system from its first training run rather than assembled afterward from separately pretrained pieces, spanning text, image, audio and video within one architecture [3].

Fact. Meta’s Chameleon pushes the same idea into generation as well as understanding. Its authors describe “early-fusion token-based mixed-modal models” that “understand and generate images and text in any arbitrary sequence,” and are explicit that reaching this required “a stable training approach from inception” — an admission that early fusion is harder to train, not merely more expensive [4].

Fact. The clearest 2026 evidence that native multimodal pretraining is now a studied, scaling-law-governed discipline rather than a one-off engineering feat comes from a July 2026 paper on scaling native pretraining from scratch, which reports that the right allocation between model scale and data mix is “highly sensitive to” data composition — and states in the same breath that training a model from scratch on large-scale multimodal corpora “typically requires expensive training costs” relative to the modular alternative [5]. The same paper is evidence for both poles of Axis A: native pretraining is real and improving, and it is not currently the cheap option.

A machined output head riding a curved rail on an any-to-any generation demonstration bench, caught mid-swing between a small horn speaker station and a resin print-bed station, its docking collar still open at neither
Figure 1. One head reaching every station on the bench is the whole ambition of an any-to-any model; here it is still between two of them, docked at neither.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Fact. The modular alternative persists for reasons a widely cited survey of efficient multimodal LLM design states directly: adaptation-centric designs — a frozen backbone plus a lightweight, separately trained connector — exist because joint-from-scratch multimodal training is costly, and because a modular system lets an operator swap in a specialist encoder for a resource-constrained or situation-specific deployment rather than retraining the whole system [6]. This is not a claim that the modular pattern is inferior; it is a claim about what it optimises for, and it optimises for something native training does not.

A wall rack of differently shaped per-modality adapter modules seated in their mounts, beside a single sealed native processing module caught mid-insertion into its own mount, its edge connector not yet fully home
Figure 2. A rack of specialists on one side, a single sealed module still going into its own slot on the other; whichever wins this rack has not decided yet.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Fact. On the perception side, Google DeepMind’s Gemini Robotics — a vision-language-action model built on Gemini 2.0 — is reported to learn new short-horizon manipulation tasks from as few as 100 demonstrations, produce smooth and reactive movement on complex manipulation tasks, and adapt to robot embodiments it was not trained on, with DeepMind’s own account stating it “more than doubles performance on a comprehensive generalization benchmark” against other vision-language-action models — a vendor claim, reported here as one [7].

ADVERTISEMENT

Fact. Physical Intelligence’s π₀ takes a related approach, pairing a pretrained vision-language model with flow matching to generate continuous robot actions, pretrained on more than 10,000 hours of data across seven robot platforms and 68 tasks, and reported performing complex tasks such as laundry folding and box assembly after fine-tuning [8].

Fact. Set against both is the most direct test of whether these results generalize. A December 2025 benchmark built specifically to test cross-domain generality evaluated GPT-5, π₀ and Magma across six capability regimes and reported that “no model demonstrates consistent generality. All exhibit substantial degradation on unseen domains, unfamiliar modalities, or cross domain task shifts,” naming “modality misalignment, output format instability, and catastrophic knowledge degradation under domain transfer” as the specific failure modes [9]. Curated-setting competence and open-world reliability are measured separately here for a reason: the same systems that clear the first do not automatically clear the second.

A compact tabletop robot arm with an overhead camera and microphone mount reaching over an open, cluttered workbench toward an unfamiliar object, its rubber-tipped gripper fingers caught open just short of closing
Figure 3. The curated stage behind it sits empty and lit; the arm itself is out on the open bench, reaching for something it was not shown in rehearsal.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Fact. On the generation side, the benchmark built specifically to test whether video-generation output is faithful to physics and commonsense — not merely visually plausible — reports that current models, even as they improve, “still struggle to generate videos that are not just visually plausible but fundamentally realistic” [10].

Fact. A companion benchmark focused on physical commonsense in action-centric video puts a number on the same gap: the best-performing model tested clears joint semantic-and-physics adherence on fewer than one in four of the benchmark’s hard test cases, and models “particularly struggle with conservation laws like mass and momentum” — the kind of failure a viewer notices without any physics background [11].

Fact. OpenAI’s own system card for Sora documents the same category of limitation from the vendor’s side: the model “may struggle with accurately simulating the physics of a complex scene” and “may not understand specific instances of cause and effect,” with objects that can vanish, deform, replicate, or pass through one another, and with output length capped in a range where quality degrades as duration grows [2].

A brass loupe on an articulated arm lowering toward a visible seam on a freshly resin-printed object on a turntable, not yet touching, beside a monitor holding a paused video frame with a visible doubling artifact at its edge
Figure 4. The flaw is already visible before the loupe ever comes down on it; checking is what turns a visible flaw into a documented one.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Fact. The push toward one model doing both understanding and generation, across any modality, is now a distinct and actively surveyed research direction rather than a single lab’s side project: a 2025–2026 survey catalogues a fast-growing set of unified architectures — diffusion-based, autoregressive-based and hybrid — built specifically to replace the pattern of a separate understanding model and a separate generation model with one that does both [12].

ADVERTISEMENT

Fact. And the field’s own evaluation infrastructure for one of the newer generative modalities, 3D content, is still being built rather than mature: the authors of the largest published text-to-3D quality-assessment database state plainly that, given “the significant quality discrepancies among various text-to-3D assets,” the field lacked “quality assessment models aligned with human subjective judgments” before their own 2025 database and model existed [13]. An evaluation infrastructure built this recently is itself evidence that production-grade quality is not yet something the field can assume.

Why these two axes, and why market structure is a consequence, not a third

Axis A is consequential because its two poles decide what kind of product a multimodal system actually is: one model trained on every modality together from the start, the way GPT-4o [1], Gemini [3] and Chameleon’s early-fusion recipe [4] already demonstrate is buildable — or a load-bearing frozen backbone with modality-specific adapters attached, kept modular specifically because joint-from-scratch training is markedly more expensive and harder to update piece by piece [6]. It is uncertain because the very paper that shows native pretraining scales cleanly also shows it costs more [5] — the two poles are not a technology waiting to be discovered so much as a cost-versus-integration trade every vendor keeps re-striking as compute prices move.

Axis B is consequential because it decides whether “multimodal AI” names something an operator can deploy without a human backstop, or something that still needs one. It is uncertain because the 2025–2026 evidence pulls in both directions inside the same eighteen months: Gemini Robotics and π₀ both demonstrate real capability inside their own evaluation settings [7, 8], while a benchmark built specifically to test generalization across those settings found that no tested model held up consistently outside them [9]; VBench-2.0’s own authors note that recent video models “perform increasingly well” on their faithfulness metrics even as they “still struggle to generate videos that are…fundamentally realistic” [10].

Whether the multimodal AI market consolidates around a small number of vertically integrated any-to-any platforms, or remains a supply chain of composed specialist vendors the way it mostly is today, is not a third axis; it is what the other two jointly produce. Write U(t)U(t) for the share of new production multimodal deployments built on a single native any-to-any model rather than a composed pipeline — Axis A’s own proxy — and R(t)R(t) for the share of multimodal deployments operating in open, uncurated conditions that meet a reliability bar without a human fallback — Axis B’s own proxy. A platform bet on one any-to-any vendor only pays off if that vendor’s model is both the one everyone is building on and reliable enough to run without a safety net; a highly reliable system stitched together from several vendors’ best components does not consolidate the market around any one of them. Consolidation is therefore better modelled as a conjunction than an average:

C(t)=1 ⁣[U(t)u]1 ⁣[R(t)r] C(t) = \mathbb{1}\!\left[U(t) \ge u^{*}\right] \cdot \mathbb{1}\!\left[R(t) \ge r^{*}\right]

C(t)C(t) stays at zero however high either term climbs alone. Axis A determines whether U(t)U(t) can plausibly clear uu^{*} within this article’s horizon; Axis B determines whether R(t)R(t) can. Neither can be inferred from the other, which is why they are kept as two axes rather than folded into one.

Four scenarios toward 2035

Crossing the two axes gives four cells. None is named as the likely outcome; each is named for the object in this article’s own room that best captures its logic.

A brushed-steel instrument cart caught mid-crossing a brass threshold strip on the floor, its rear wheels still on a spotlit curated demonstration alcove and its front wheels already on the open, cluttered workbench floor beyond
Figure 5. Curated and open sit right next to each other in this room; the distance between them is exactly one cart's length, and this cart is only halfway across it.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Scenario one: The Docking Head — architecture converges, reliability arrives

Mechanism. Native any-to-any training becomes the default recipe as the cost gap the native-scaling paper documents narrows [5], and the resulting single models close the generalization gap MultiNet’s benchmark measured in 2025 [9], so that one model reliably serves both perception — real-time understanding good enough for embodied robotics, extending the trajectory Gemini Robotics and π₀ are already on [7, 8] — and generation — video and 3D output good enough to ship without a human pass, extending the trajectory VBench-2.0 and VideoPhy-2 track [10, 11] — without a specialist pipeline behind it.

Horizon. Recognisable architectural convergence by 2030–2031; a reliable, dominant any-to-any deployment pattern plausible by 2034–2035.

Assumptions. The cost gap the native-scaling paper documents keeps narrowing rather than persisting as a structural tax; the generalization gap MultiNet measured in December 2025 closes rather than proving durable.

Observable indicators. A single native model becomes the default backend behind commercial embodied-robotics deployments and commercial video or 3D generation products, replacing rather than sitting alongside per-modality specialist services; frontier system cards report open-world success rates approaching curated-benchmark rates rather than trailing them by a wide margin.

Disconfirmation. Falsified if, by 2032, the leading commercial deployments in both robotics and generation are still built on composed specialist pipelines rather than one native model, or if published open-world success rates remain far below curated-benchmark rates.

Scenario two: The Module on the Stage — architecture converges, reliability stalls

Mechanism. Architecture consolidates around native any-to-any models — the GPT-4o, Gemini and Chameleon trajectory becomes the industry-standard training recipe [1, 3, 4] — but reliability does not close at the same pace. The generalization failures MultiNet documents and the physics and commonsense gaps VideoPhy-2 and the Sora system card document turn out to be about training-data coverage and evaluation reach rather than architecture, so a single elegant model still performs unevenly outside curated settings [9, 11, 2].

Horizon. Architectural convergence recognisable by 2029–2030; a reliability gap that persists through 2035 without necessarily closing.

Assumptions. Architecture and reliability are more separable than Axis A’s optimistic pole assumes; investment concentrates on unifying model families faster than on the harder problem of open-world data coverage.

Observable indicators. Native any-to-any models become the default even for narrow production tasks, yet system cards keep reporting large curated-versus-open-world performance gaps; specialist components persist specifically as a reliability patch — a fallback verifier, a human-in-the-loop wrapper — bolted onto an otherwise unified model rather than disappearing.

Disconfirmation. Falsified if open-world reliability indicators converge toward curated-benchmark indicators within the horizon — that would indicate the world has moved to scenario one instead.

Scenario three: The Rack on the Open Floor — architecture stays modular, reliability arrives

Mechanism. The reliability gap closes through engineering rather than unification. Composed pipelines of specialist, independently improved components — a dedicated perception encoder, a dedicated action model, a dedicated generation model, orchestrated together — reach open-world dependability the way narrowly scoped systems more easily do, consistent with the efficient-MLLM literature’s own rationale for keeping components swappable and independently upgradable [6], while no single native model ever becomes the default backend.

Horizon. Recognisable by 2029–2030 as reliability metrics on composed pipelines pull ahead of native-model metrics; stable through 2035 if no single model closes the same gap on its own.

Assumptions. Reliability is fundamentally a data-coverage and specialization problem rather than a fusion-architecture problem, so a narrowly scoped specialist keeps beating a generalist asked to cover six capability regimes at once — the reading MultiNet’s own degradation finding supports [9]; orchestration overhead stays manageable at scale.

Observable indicators. Production embodied-robotics and generation deployments report open-world reliability at or near curated-benchmark levels while remaining architecturally composed — multiple named vendor components in one pipeline — rather than backed by one native model; no vendor’s any-to-any model displaces a competitor’s best-in-class specialist product for a given modality.

Disconfirmation. Falsified if a single native any-to-any model becomes the reported backend behind the majority of new reliable deployments — that would indicate the world has moved to scenario one instead.

Scenario four: The Long Rehearsal — architecture stays modular, reliability stalls

Mechanism. Neither pressure resolves. Composed, adapter-based pipelines remain the pragmatic default for the cost and flexibility reasons already documented [6], and the reliability gap the current evidence already shows — MultiNet’s generalization failures, VideoPhy-2’s ceiling on the hard subset, the Sora system card’s physics and causality caveats — continues largely unresolved [9, 11, 2]. This scenario’s leading indicator is not a future event; it is the documented present above, continuing.

Horizon. Close to today’s baseline; recognisable as the stable case by 2028 if none of the trends above accelerate; could persist through 2035.

Assumptions. No single architectural or data breakthrough forces convergence on either axis; competitive incentives keep favouring incremental, composable improvement over a costly platform bet.

Observable indicators. Published reliability figures stay within roughly today’s range through the horizon; specialist per-modality vendors continue to compete effectively against any-to-any platforms on narrow benchmarks; curated demonstrations continue to substantially outperform open-world deployment in system cards.

Disconfirmation. Falsified if either architectural convergence or open-world reliability is observed at the thresholds defined in scenarios one through three — either observation would move the world out of this cell.

What all four share, and the wildcard neither axis names

Three things hold across every cell. First, a human-in-the-loop or human-reviewed fallback survives in three of the four scenarios and only fully disappears in the one — The Docking Head — where both axes resolve toward the optimistic pole simultaneously; even Gemini Robotics and π₀, the two systems furthest along today, are demonstrated and reported inside evaluation settings a person designed and is watching [7, 8]. Second, which specific fusion recipe wins inside Axis A’s convergent pole — Chameleon-style early fusion, a GPT-4o-style end-to-end recipe, or a successor neither has published yet — is a narrower, more replaceable engineering detail than the axis itself; the axis is compatible with any winner. Third, none of the four requires a plateau in the surrounding capability trend that makes the question worth asking in the first place. A model can keep improving on held-out text and reasoning benchmarks while multimodal reliability stays exactly where MultiNet and VideoPhy-2 found it, because these axes describe multimodal maturity specifically, not general capability growth.

The four scenarios share a blind spot too. All four assume the roughly gradual movement the 2025–2026 evidence already shows. A large enough shock would not fit cleanly into any of them. A single training or data technique that closed the generalization gap MultiNet measured — the way weak supervision at scale closed earlier gaps in single-modality systems — could move Axis B abruptly rather than along any of the four scenarios’ gradual paths, and drag Axis A’s economics with it, since a reliable native model suddenly becomes worth its training-cost premium. Equally, a widely publicised embodied-robotics or generated-media failure — of the kind the Sora system card and the VideoPhy-2 benchmark already document in miniature — landing inside a high-stakes deployment rather than a benchmark could set trust back further and faster than The Long Rehearsal’s gradual baseline implies. Either event would move both axes at once, abruptly, rather than along the paths each scenario traces on its own.

Two predictions, stated separately from the scenarios

Prediction one. Horizon: end of 2029. At least one frontier lab’s flagship multimodal system will report open-world or uncurated evaluation results explicitly alongside its curated benchmark results, with the gap between the two reported as a named metric rather than left implicit. Assumption: the generalization-gap finding MultiNet reported in December 2025 becomes something vendors are pressured to address rather than something the field quietly stops measuring [9]. Indicator: a system card, technical report, or benchmark submission that reports a curated-versus-open-world or in-distribution-versus-out-of-distribution delta as a headline number rather than a footnote. Disconfirmed if leading multimodal system cards in 2029 still report only curated-benchmark scores with no open-world or generalization delta disclosed.

Prediction two. Horizon: end of 2031. No single vendor’s any-to-any model will have become the reported backend behind a majority of production embodied-robotics deployments; composed pipelines with a dedicated perception or action component will remain more common in production than a single native model handling perception, planning and action together. Assumption: the cost and data advantages of narrow specialization documented in the efficient-MLLM literature continue to outweigh the integration benefits of a single native model for safety-critical physical deployments specifically, even where they do not for lower-stakes consumer applications [6]. Indicator: vendor deployment disclosures, integrator case studies, or procurement documentation naming the components of a production robotics stack. Disconfirmed if a majority of newly deployed production robotics systems by the horizon are documented as running on one vendor’s native any-to-any model rather than a composed stack.

Neither prediction requires a capability discontinuity. Both follow from the structure already visible: two genuinely uncertain axes, a present documented well enough to anchor a horizon, and a field that has only recently started reporting the specific numbers — an open-world delta, a cross-domain degradation rate — that would let anyone outside the labs tell which cell of the matrix the world is actually in.

What to take away

The refusal to pick a favourite among these four is the substantive claim, not a hedge around one. In mid-2026 the evidence is genuinely split on both axes at once: a model trained end-to-end across text, vision and audio from a single vendor sits beside a survey explaining, in the same season, why adapter-based modularity remains the pragmatic default for everyone without that vendor’s training budget [1, 6]. A robotics model that learns a new task from a hundred demonstrations sits beside a benchmark built specifically to show that the same class of model degrades sharply the moment the task moves outside its training distribution [7, 9]. A video-generation benchmark suite built to track physics faithfulness sits beside its own finding that the best current model still fails the hard physics cases more often than it passes them [10, 11]. Anyone reporting a single confident future for multimodal AI in 2035 is reporting which of these four scenarios they would bet on, not what the current record shows.

The more useful discipline is the one this article tried to practise: name the axis, ground each pole in a dated result, and say in advance what observation would mean the world had moved to a different cell. Multimodal AI has already demonstrated real things — native training at scale, real-time robot control from language, generation good enough to fool a casual viewer. Whether any of that generalizes into the reliable, unified, everywhere infrastructure its more confident advocates already describe it as becoming is still open.