A Sweet Potato, Two Home Videos, and a Robot Nobody Had Trained to Cook
On April 16, 2026, Physical Intelligence told reporters its newest robot policy had done something none of its earlier models had: cook a sweet potato in a household air fryer it had never been trained to operate [9]. The company’s own account of what made this possible is unusually specific for an industry that mostly deals in adjectives. The robot’s training data included exactly two home-recorded episodes involving that particular appliance — one labeled “push the frying basket into the airfryer,” the other “put the basket of the airfryer on the leftmost side of the counter” — plus unrelated footage of a completely different robot, a Franka arm, drawn from the public DROID dataset [8] [9]. Given nothing but a plain-language instruction and zero task-specific coaching, the model’s first zero-shot attempt was, in TechCrunch’s own description, “a passable attempt.” In a separate account from the same report, a Physical Intelligence researcher describes an earlier version of the same test producing a 5 percent success rate that reached 95 percent after about thirty minutes of a human refining how the task was described in language [9]. The company calls the underlying policy, π0.7, a “Steerable Model with Emergent Capabilities” and says it exhibits “a step-change in generalization” [8].
That is a real result, reported with more precision about its own inputs than almost anything else this piece reviewed. It is also, on its own, a single demonstration — and a single demonstration cannot answer the question its headline implies. NVIDIA says its GR00T N1 foundation model lets a humanoid “reason about novel situations” and “robustly handle real-world variability” [1]. Figure says its Helix model can “pick up virtually any small household object, including thousands of items they have never encountered before,” and even generalizes to abstract categories — asked to “pick up the desert item,” Helix identifies a toy cactus as the match and executes the grasp [2]. Three companies, three anecdotes, and not one shared unit for how far outside each model’s own training distribution the anecdote in question actually reaches. “Generalizes” is doing enormous work in 2026 humanoid coverage, and almost nobody saying it has defined what would make the claim false.
The scale of the problem this creates is not a matter of tone. A June 22, 2026 meta-analysis of 1,228 vision-language-action (VLA) papers, spanning arXiv postings from February 2023 through June 2026, cites a May 2026 diagnostic study finding that collapsing manipulation competence into a single binary success rate can inflate reported performance by up to 70 percent relative to a more fine-grained accounting of what actually happened during a trial [10] [13]. If the field’s own standard reporting unit already overstates success by that much, a marketing anecdote sitting one further step removed from any benchmark at all is not a rounding error on top of an otherwise reliable number. It is evidence for a different claim entirely, dressed in the same word.
An Anecdote Reports a Success. It Does Not Report a Distance.
Here is the distinction the marketing keeps blurring. A demonstration — the cactus, the air fryer, GR00T N1’s “novel situation” — reports that a policy succeeded once, on one instance, under conditions a reader is invited to imagine as difficult. A systematic evaluation reports, in a methods section, which specific dimensions of the training distribution were deliberately varied, and by how much, before the model was scored. These are not the same kind of evidence, and treating a well-produced anecdote as a substitute for the second thing is the exact substitution I want to name and then refuse for the rest of this piece.
There are four dimensions along which a robot policy’s training distribution can meaningfully stretch, and this piece scores all four lineages against only these: embodiment (does the evaluation run on robot hardware, a gripper, or a body plan the policy was not trained on), environment (does it run in a physical space, lighting condition, or clutter pattern outside the training set), object category (does it manipulate an item, or a class of items, absent from training), and instruction phrasing (does it correctly execute a command worded, or composed, in a way training data did not cover). A single anecdote can gesture at all four axes moving at once, or at none, and a reader watching a thirty-second demo video has no way to tell which. A disclosed methodology section can state, in plain language, exactly which of the four moved and by how much — and, notably, several of the makers surveyed here already do this, at least some of the time, which is the strongest evidence in this piece that the standard is achievable rather than aspirational.
I call the count of axes a maker’s own disclosed evaluation explicitly and systematically varies its generalization radius for that release. It runs from zero (the evaluation tests only what training already covered) to four (all of embodiment, environment, object category, and instruction phrasing are deliberately varied and the variation is disclosed with enough specificity that an outside reader could, in principle, check it). Radius is scored only from a maker’s own public statements about its own evaluation — never from third-party benchmark leaderboards, never from this piece’s own guesses about what a company probably tested but did not publish. Where a maker discloses nothing about an axis, that axis scores zero for that release, even if the company privately tested it and simply did not write it up. That asymmetry is a real limitation of the method, not a footnote to wave away: a company with a thin public blog post and rigorous internal evaluation will score worse here than a company with a thin evaluation and an eloquent one. What the method still buys, even carrying that limitation, is a same-source time series — tracking one company’s own disclosure practice across its own dated releases holds the company’s incentives and writing style roughly constant, which a single cross-company snapshot cannot do. That is why the richest part of this piece is not the three-way comparison across NVIDIA, Figure, and Physical Intelligence. It is Physical Intelligence’s own four dated releases, read against each other.
NVIDIA’s Own Paper Names an Axis and Leaves the Count Off the Page
GR00T N1, published by a 41-author NVIDIA team on March 18, 2025 and lightly revised nine days later, is built as a dual-system architecture: a vision-language module the paper calls System 2, running slower and handling scene understanding and instruction interpretation, paired with a diffusion-transformer module called System 1 that generates motor actions in real time [1]. The model was trained on what the paper describes as “a heterogeneous mixture of real-robot trajectories, human videos, and synthetically generated datasets,” tested across multiple robot embodiments in simulation benchmarks, and deployed for language-conditioned bimanual manipulation on a Fourier GR-1 humanoid [1].
Score that disclosure against the four axes and one clears the bar: embodiment, since the paper explicitly states the model was evaluated across more than one robot body and names a specific humanoid deployment. Even that one axis is unquantified — the paper does not state how many embodiments “multiple” means, nor how different they are from one another along any measurable dimension. Environment, object category, and instruction phrasing are not broken out with specific evaluation numbers anywhere in the disclosed materials this piece could verify; the abstract’s own language — “novel situations,” “real-world variability,” “rapidly learn new tasks” — implies motion along all three, but implies rather than states. Disclosed radius for GR00T N1: one axis, and that one unquantified. Anecdote-implied scope, reading the abstract’s own claims at face value: all four. This is the widest gap this piece found between what a lineage’s marketing language implies and what its own technical disclosure operationalizes — not necessarily because GR00T N1 generalizes less well than its rivals, but because its public accounting of its own evaluation is the least specific among the four lineages surveyed here.
Figure’s Helix Is the One Lineage Whose Anecdote and Disclosure Roughly Agree
Figure’s Helix, announced February 20, 2025, is a “System 1, System 2” vision-language-action model in the same broad family as GR00T N1 — a vision-language component running at 7–9 Hz for scene and instruction understanding, paired with a fast reactive policy running at 200 Hz for motor control [2]. What sets Helix’s launch materials apart from GR00T N1’s is specificity at the object-category axis: Figure states its robots handled “thousands of novel items in clutter — from glassware and toys to tools and clothing — without any prior demonstrations or custom programming,” in what the company calls systematic testing rather than a single cherry-picked run [2]. That is an object-category claim disclosed with an actual quantity and named subcategories, not a single held-out example standing in for the whole axis.
Helix’s flagship anecdote also does something the GR00T N1 abstract does not: it names a second axis explicitly, rather than only implying it. Asked to “pick up the desert item,” Helix identifies a toy cactus as the match and executes the grasp — a demonstration Figure frames not as object recognition but as abstract-concept mapping, selecting “the closest hand” and computing “the precise motor commands needed to grasp it securely” once the correct referent is identified [2]. That is a real second axis, instruction phrasing at a level of abstraction beyond a literal object name, and unlike GR00T N1’s broad abstract language, Helix’s own anecdote and its own quantified category claim describe the same two axes rather than one gesturing vaguely at all four. Disclosed radius for Helix’s 2025 launch: two axes, both with real specificity behind them — object category, quantified in the thousands and by named subcategory, and instruction phrasing, demonstrated through abstract-concept resolution rather than literal matching. This is the narrowest gap between anecdote and disclosure this piece found among the three non-Physical-Intelligence lineages.
What Helix’s own materials do not cover is just as informative. There is no disclosed claim about the environment axis — Figure’s deployments run in factories and warehouses the company itself operates and instruments, not in homes or spaces the model had never encountered, the way Physical Intelligence’s π0.5 explicitly tests below. There is no disclosed cross-embodiment claim; Helix runs on Figure’s own hardware only. By the time of Figure’s June 30, 2026 update, the axis Figure chooses to talk about shifts again: Helix 02, the version now deployed at BMW’s Spartanburg plant, is described as coordinating “Figure 03’s hands, arms, torso, and feet” through “pixels-to-actions” control on a live automotive logistics line [3]. That is a genuine and difficult engineering claim — whole-body sensorimotor coordination under real production constraints — but it is a claim about control-loop integration, not about which training-distribution axes the underlying policy has learned to cross. Reading it as a generalization-radius update rather than what it actually is would be exactly the kind of category error this piece exists to catch.
Physical Intelligence’s Four Dated Releases Are the Only Time Series in the Survey
NVIDIA and Figure each supplied one dated snapshot to score. Physical Intelligence supplied four, released over eighteen months and independently datable against its own blog index [4], which makes its own publication history the only place in this survey where generalization radius can be tracked release over release from a single, self-consistent source rather than compared once across companies with different incentives and different house styles for what to disclose.
π0, announced October 31, 2024, is the baseline. Physical Intelligence trained the policy on eight distinct robots and evaluated it on five tasks spanning three separate hardware platforms — a UR5e arm, a bimanual ARX platform, and a bimanual Trossen platform [5]. That is the embodiment axis, disclosed with an actual, checkable roster rather than a vague “multiple robots” claim, and it is worth noting that here the company’s own marketing language undershoots its technical disclosure rather than the more common pattern of overshooting it: the launch post’s flagship phrasing — a policy that can “perform a wide range of different skills and control a wide range of different robots” — is vaguer than the specific eight-robot, five-task roster sitting one paragraph away in the same post. Disclosed radius for π0: one axis, embodiment, and this is the one release in the survey where the anecdote is less ambitious than the disclosure behind it.
π0.5, announced April 22, 2025, is the strongest multi-axis disclosure this piece found anywhere. Physical Intelligence put the policy into “entirely new homes” it had never operated in and reported the data curve behind that claim rather than only the headline result: performance climbed with more than 100 distinct training environments, and the smallest documented ablation used roughly 400 hours of mobile-manipulation data [6]. That is the environment axis, quantified twice over — once by environment count, once by data volume. The same release states its flagship tasks were driven by single high-level natural-language commands — “clean the bedroom,” “put the dishes in the sink” — standing in for entire unscripted multi-step sequences the policy had to plan out on its own, including behaviors like using a sponge to address a spill that the instruction never spelled out [6]. That is the instruction-phrasing axis, and it is disclosed as compositional planning rather than literal command-matching. Physical Intelligence is also candid about the limit: the post states plainly that the policy “does not always succeed on the first try” [6], which is more honesty about failure than either of the other two companies’ launch materials volunteer. Disclosed radius for π0.5: two confirmed axes, environment and instruction, both quantified with real numbers rather than adjectives — the single most operationalized disclosure in this survey outside Physical Intelligence’s own later release, and closely matched by its own anecdote rather than exceeding it.
π0.6, announced November 17, 2025, is the negative case this piece needs to make the whole framework mean anything. No new axis appears. The three named tasks — making espresso, folding laundry, assembling packaging boxes — are the same tasks Physical Intelligence had already been demonstrating; what changes is how reliably the policy performs them, after reinforcement learning on autonomously collected experience. On the hardest of the three tasks, espresso-making, throughput roughly doubled, from about 15 successful drinks per hour to more than 30, while the success rate climbed from around 40 percent to over 90 percent [7]. Physical Intelligence reports the policy then ran an entire day making espresso, folded fifty different novel laundry items in a new home for hours without interruption, and assembled and labeled 59 packaging boxes in a real factory [7]. Every one of those numbers is a genuine, well-documented reliability result. None of it is a generalization-radius result. Disclosed radius for π0.6: zero new axes relative to π0.5 — the improvement lands entirely on supervision and shift-length reliability for a fixed task set, which is a different, equally real claim that a different entry in this series tracks on its own terms. A reader skimming past the fine print could easily fold π0.6’s throughput gains into the same “generalizes further” narrative used for π0.5 and π0.7. That conflation — mistaking “does the same thing more reliably” for “does more different things” — is precisely the substitution generalization radius is built to catch, and π0.6 is the cleanest example in this survey of why the two claims need to stay separated.
π0.7, announced April 16, 2026, is where the survey opened, and against its own three predecessors it is both the widest radius Physical Intelligence has disclosed and the most precisely documented. The air fryer demonstration names its own inputs — two home episodes plus unrelated Franka-arm footage from the public DROID dataset [8] [9] — which puts at least three axes in motion at once and states as much rather than leaving a reader to infer it: object/task category (an appliance the deployed robot had two loosely related examples of, not a trained-on task), embodiment-transfer (data drawn from an entirely different robot’s arm), and skill composition, which Physical Intelligence names directly as the mechanism — “recombining skills from various tasks to solve new problems.” A second demonstration makes the transfer claim explicit as a fact about the training set rather than an inference from a video: the policy folded laundry on a bimanual UR5e platform with “no training data… for this task with the UR5e” at all [8]. Disclosed radius for π0.7: three axes, each named with enough specificity — episode counts, a named source dataset, an explicit statement of zero task-specific training data for one embodiment — that an outside reader could, in principle, check the claim against Physical Intelligence’s own account rather than take the demo video’s word for it.
The Four Lineages, Scored on Their Own Disclosures
| Lineage (date) | What the disclosure names | Axes explicitly varied (of 4) | Flagship anecdote | Anecdote-to-disclosure gap |
|---|---|---|---|---|
| NVIDIA GR00T N1 (Mar 2025) [1] | “Multiple robot embodiments” in simulation; one named deployment (Fourier GR-1) | 1 (embodiment, unquantified) | “Novel situations… real-world variability… new tasks” | Widest — anecdote language implies all four axes; disclosure operationalizes one |
| Figure Helix (Feb 2025) [2] | Thousands of novel items across named categories; “desert item” → cactus abstraction | 2 (object category, instruction) | The same cactus example used as both the anecdote and the evidence | Narrow — anecdote and disclosure describe the same two axes |
| Physical Intelligence π0 (Oct 2024) [5] | 8 robots trained; 5 eval tasks across 3 hardware platforms | 1 (embodiment, quantified) | “Wide range of different skills” (qualitative) | Inverted — disclosure exceeds the anecdote’s own ambition |
| Physical Intelligence π0.5 (Apr 2025) [6] | 100+ new homes; ~400h smallest data ablation; single-command task planning | 2 (environment, instruction — both quantified) | New-home kitchen/bedroom cleanup from one command | Narrow — anecdote matches the disclosed multi-axis claim |
| Physical Intelligence π*0.6 (Nov 2025) [7] | Same 3 known tasks; RL on autonomously collected experience | 0 (a reliability axis, not one of the four) | A full unattended day of espresso-making | Not applicable — this release claims reliability, not radius |
| Physical Intelligence π0.7 (Apr 2026) [8] [9] | 2 home episodes plus cross-embodiment DROID data; laundry transferred to an untrained UR5e | 3 (object/task, embodiment-transfer, skill composition) | Sweet potato in the air fryer, presented with its own data provenance | Narrowest — the disclosure names its own inputs down to the episode count |
The Metrics-Inflation Finding Explains Why Almost Nobody Reports This
The pattern in that table — some lineages disclosing real axis counts, most not, none using a shared unit — is not an oversight the field will casually correct, because the field’s own default reporting convention has no room for the distinction in the first place. The diagnostic study behind the “up to 70 percent” inflation figure did not simply re-run existing benchmarks and find a bigger error bar. It built a framework, which its authors call MetaFine, that separates manipulation competence into three distinguishable dimensions — understanding, perception, and controlled behavior — and re-scores existing benchmark trials against all three instead of collapsing them into one pass/fail bit [10]. A single binary success rate cannot, by construction, tell a reader whether a trial succeeded because the policy solved a genuinely unfamiliar combination of skills or because the task sat one small step from ten near-identical tasks already in its training set. Both register as identical successes. That is the same collapse driving the anecdote problem this piece opened on, one level down in the stack: if the field’s own standard benchmark cannot distinguish narrow competence from real generalization at the level of a single trial, a company choosing between publishing a rigorous four-axis breakdown and publishing one impressive video has very little technical pressure pushing it toward the former.
Two further findings from the same June 2026 meta-analysis sharpen exactly where that collapse does its damage. LangGap, a February 2026 benchmark built specifically to probe the instruction-phrasing axis, found that targeted, diverse instruction augmentation — relabeling existing trajectories with more varied wording, no new physical demonstrations at all — took a single-task success rate from 0 percent to 90 percent [11] [13]. That sounds like exactly the kind of instruction-phrasing generalization this piece has been scoring as a real axis. But the same paper reports that the identical technique, extended to cover multiple tasks at once rather than one, only reached 28 percent [11] — a gap between the single-task and multi-task numbers that is itself a measurement of how much of “generalizes to new phrasing” was actually closer to memorizing more phrasings of one already-known instruction. A November 2025 data-distillation result adds a second angle on the same problem from the volume side: a curated coreset containing just 5 percent of a full VLA training dataset recovered 85–90 percent of the full dataset’s performance [12] [13]. If the overwhelming majority of a training set is not doing the work that determines final performance, then “trained on more data” and “trained on data that actually spans more of the four axes this piece is scoring” are being treated as the same claim when they plainly are not — and a headline anecdote, sitting downstream of both a binary success metric and an undifferentiated pile of training data, is the least equipped format available to tell the two apart.
What Would Make This Complaint Obsolete
The claim this piece is making is narrow enough to state a clean failure condition for it. If humanoid foundation-model makers begin disclosing generalization radius, in this explicit multi-axis form or an equivalent one, as a routine part of a release rather than as a single flagship anecdote, the complaint driving this entire piece is resolved and the diagnostic stops being useful. That is not a hypothetical future state contingent on some new industry standards body. Physical Intelligence’s own π0.5 and π0.7 releases already do most of what this piece is asking for, in public, using no more than the same blog format every other release in this survey used. The unevenness is the finding: the same company that names exact environment counts and episode counts for two of its releases discloses nothing more than a qualitative catchphrase for the other two, in a pattern that tracks what each specific release was actually claiming rather than any consistent house standard for what a launch post owes its readers.
The honest objection to all of this is the one raised earlier and worth restating plainly rather than burying: counting disclosed axes rewards whichever lab writes the most detailed blog post, not necessarily whichever model reaches furthest beyond its training data. A company could test rigorously across all four axes internally and simply not publish the breakdown, scoring zero here despite having done the work; a company with a thinner internal evaluation but a well-drafted release could score higher. That asymmetry is real, and no scoring method built only from public disclosures can fully escape it. What survives the objection is narrower than “generalization radius measures true generalization” — it is that generalization radius measures what a maker is willing to state, checkably, about its own evaluation, and that this number is worth more to an outside reader than an anecdote precisely because it can be wrong in a nameable way, checked against a specific published claim, in a way a video of a robot cooking a sweet potato cannot.
That distinction has a direct use for anyone whose decision does not get to wait for the field to fix its own reporting habits. An integrator evaluating a VLA stack for a task the vendor has not specifically demoed — an unfamiliar SKU next quarter, a warehouse the model has never operated in — is better served asking which of the four axes the vendor’s own evaluation ever varied, and by how much, than asking how convincing the vendor’s demo video looks. NVIDIA’s own paper, on the numbers it has chosen to publish, currently answers that question for one axis. Figure’s answers it for two, matched closely to its own flagship claim. Physical Intelligence answers it for three, in its most recent and most specific release, and for zero, honestly, in the release immediately before it, where the actual claim being made was about reliability rather than reach. Until the rest of the industry starts publishing something closer to that breakdown as a matter of course, the number worth asking after when a new humanoid foundation model claims it generalizes is not how many degrees of freedom the arm has or how good the demo footage looks. It is how many of four specific things actually moved, and whether the company saying so is willing to name the count.