Humanoid Foundation Models All Claim to Generalize: Almost None Say How Far
NVIDIA, Figure, and Physical Intelligence each market a flagship anecdote as proof their robot brain generalizes. None report, in a shared unit
Tagged foundation models · show all articles
NVIDIA, Figure, and Physical Intelligence each market a flagship anecdote as proof their robot brain generalizes. None report, in a shared unit
A token has no modality, but a bill does. Priced out per image, per second of video, per training pair and per byte moved, multimodal AI turns out to be several different economic activities wearing one name.
Leaderboards read like steady progress toward one unified system. Five specific capabilities are not solved, by the researchers' own published admission — and each has a benchmark built to prove it.
A model can post a state-of-the-art score on a benchmark built to require sight without once needing to look. Modality ablation is the test that tells the difference between what a system saw and what it merely guessed.
A model that sees, hears and reads still returns one fluent answer, and fluency hides the fact that the link between that answer and what was actually there breaks in ten distinct, independently documented ways.
Two questions decide how multimodal AI matures by 2035: whether architecture converges on native any-to-any models, and whether perception and generation both cross into open-world reliability. Four scenarios, each with a falsifier.
Image-text systems tokenize a grid and align two vectors. Video is a spacetime volume, audio is a waveform, and a 3D scene is a continuous field — none of them arrives as a grid, and each forces its own discretization scheme before any model can read it.
Four research lineages solved the same problem — joining vision, audio and text into one system — with four different training recipes. Each recipe holds one cost fixed and lets another float, and the documented trade is an engineering choice, not a hierarchy.
Before a single model could read an image and write about it, two research communities spent a decade building separate pieces that had to be joined by hand. This is the record of each join.
Claude carries more defensive machinery than most deployed models — classifiers, citations, constitutional training. Its failures surface at the seams of that machinery, and Anthropic's own records are usually the first place they are documented.