An intention becomes torque

A trained robot policy is, structurally, nothing more than a function: an observation goes in, an action comes out. Explaining why that function is hard to learn — the data that has to be produced by the machine being built, the parts of physics a simulator does not model, the deadline a control loop cannot miss — is a separate piece of work, already done elsewhere in this series. This article does not repeat that argument. It opens the box the previous one left closed and asks a narrower, more mechanical question: once a policy exists, what actually happens between a camera frame arriving and a motor winding drawing current?

The honest answer is a pipeline with five distinct stages, each running at its own frequency, trained by at least three different methods that get combined rather than chosen between, rehearsed in simulation before any of it touches a real joint, and finally handed to a layer of classical control that is allowed to override the learned part entirely. None of that is exotic. It is closer to how a modern car’s driver-assistance stack is built than to how a language model answers a question, and understanding it means walking the loop in order.

The loop underneath the policy

Every embodied policy, learned or classical, sits inside the same five-stage loop: sensing, state estimation, planning or policy inference, control, and actuation. The stages are not a metaphor for a pipeline diagram; they are separate pieces of hardware and software running at different rates, and the rate mismatch between them is itself an engineering problem, not an afterthought.

ADVERTISEMENT

Sensing supplies raw signal: RGB and depth images from wrist- and scene-mounted cameras, joint encoder counts, six-axis force-torque readings at a wrist or foot, and on legged platforms an inertial measurement unit reporting orientation and angular rate. None of this is directly usable on its own. Cameras report pixels, not the pose of a graspable object; encoders report angles, not where a foot actually is relative to the ground it is about to strike.

State estimation is the stage that turns raw signal into a state a controller can act on: filtering noisy proprioception, fusing exteroceptive and proprioceptive streams, and producing an estimate of pose, velocity and contact status that is consistent from one instant to the next. Miki and colleagues built this fusion explicitly rather than by hand-tuned heuristic, training an attention-based recurrent encoder end to end to combine a quadruped’s exteroceptive terrain sensing with its proprioceptive body feedback, specifically because the two disagree in ordinary conditions — snow, tall grass and standing water all look like solid ground or like obstacles to a camera in ways that proprioception alone can correct for [9]. State estimation typically runs fast and continuously, because everything downstream depends on it being current.

A stereo camera and depth sensor on a calibration gantry above a checkerboard target, a thin laser alignment line crossing the board not yet centred on the target's origin mark
Figure 1. State estimation starts with knowing exactly where a sensor is; a policy trained on unaligned geometry is trained on a world that is not quite the one the robot occupies.

Only after a state estimate exists does a policy — learned or scripted — decide what to do next. This is the stage that receives the most public attention and runs the slowest of the decision-adjacent layers. A vision-language-action model has to encode an image and a language instruction through a large pretrained backbone before it can emit anything, and that cost is real: pi-zero, built as a continuous-action head on top of a pretrained vision-language backbone, is documented as producing actions at up to 50 Hz for dexterous tasks, and it emits them in chunks of many future actions per forward pass specifically so the next expensive forward pass does not have to complete before the arm can keep moving [2]. Fifty hertz sounds fast in conversation and is slow by the standard of the stage beneath it.

Control is where a policy’s output — a target end-effector pose, a target joint configuration, a desired body velocity — is converted into a torque or current command that respects the robot’s actual dynamics, its joint limits, and, for a legged or manipulating system in contact, the forces at that contact. This stage runs an order of magnitude or more faster than policy inference, because it is what keeps a limb from overshooting a target or a leg from buckling in the interval between two policy decisions. Actuation is the last stage and the one every other stage exists to serve: current through a winding becomes torque at a joint, and the loop closes when that torque changes what the sensors report next.

Treat the loop as one object and the design problem becomes visible. A slow-reasoning policy can still control a robot well if the layer beneath it holds position and rejects disturbance during the gap. A fast-reasoning policy on a slow, poorly tuned control layer cannot.

ADVERTISEMENT

Teaching a policy by demonstration

The most direct way to install a policy is to show it what to do and have it copy the demonstration — imitation learning, and specifically behavior cloning: collect pairs of observations and the actions a human took in that observation, and train a network to reproduce the mapping. The demonstrations overwhelmingly come from teleoperation, because a manipulation trajectory is a record of a specific body moving through a specific controller, and there is no corpus of that lying around to scrape. A human operator drives a leader arm or a handheld interface; a follower arm, or the same robot under a different control mode, reproduces the motion while every joint command and every camera frame is logged. The scale this actually requires is worth stating in concrete terms: DROID, one purpose-built collection effort, reports 76,000 demonstration trajectories amounting to 350 hours of interaction, gathered across 564 scenes and 84 tasks by 50 human collectors working across three continents over twelve months [6]. That is the labour cost of one dataset, for one slice of the manipulation task space.

A teleoperation leader arm mid-gesture on a lab bench, its follower arm's gripper closing on a small part with a status indicator lit on the recording unit beside it
Figure 2. Every demonstration a policy imitates was moved through by a human operator first; the recording captures the motion, not the intention behind it.

What a network trained on that data actually predicts is not obvious, and the choice of action representation has turned out to matter as much as the choice of backbone. Naive behavior cloning predicts one action per timestep and tends to average across the different ways a demonstrator might have completed a task, producing a blurred, indecisive policy at exactly the moments where a task has more than one valid solution — approaching a mug from the left or the right, say. Two responses to that problem now dominate the field’s practice, and they came from different directions.

Zhao and colleagues introduced Action Chunking with Transformers specifically to combat compounding error in imitation learning: rather than predict one action, the policy predicts a whole chunk of future actions in a single forward pass and executes them together, which reduces both how often the network is queried and the number of places an error can be reintroduced mid-task. On six fine manipulation tasks — including opening a translucent condiment cup and inserting a battery — the method reached 80 to 90 percent success from around ten minutes of demonstrations per task, on hardware the authors built to be inexpensive to reproduce [4].

Chi and colleagues took a different route to the same problem, representing the policy itself as a conditional denoising diffusion process over action sequences rather than a direct regression. Because a diffusion model can represent an arbitrarily multimodal distribution rather than committing to one mean, it handles exactly the branching-choice cases that blur a naive policy; combined with receding-horizon execution and visual conditioning, it produced an average 46.9 percent improvement over prior methods across fifteen tasks drawn from four manipulation benchmarks [5].

Neither method changes what a demonstration fundamentally is: a recording of one solution to a task, produced by a human, that a network is asked to generalise from. What they change is how much of that recording’s structure — the shape of a whole gesture rather than a single instant, the existence of more than one right answer — survives into the trained policy.

The vision-language-action recipe

Imitation learning explains how a policy is fit to data. It does not explain where a model’s ability to generalise beyond exactly what it saw comes from, and the answer the field has converged on is to not start from a blank network at all. A vision-language-action (VLA) model takes a vision-language backbone that has already been pretrained on internet-scale image-and-text data — so it already has some working notion of what a mug, a drawer or the word “leftmost” refers to — and bolts an action-producing head onto it, then trains the whole thing further on robot trajectories.

ADVERTISEMENT

The two dominant ways of building that action head illustrate a real design choice rather than a cosmetic difference. RT-2 represents robot actions as tokens drawn from the same discrete vocabulary the underlying language model already uses for text, so producing an action is, mechanically, the same operation as producing the next word in a sentence; the model is co-fine-tuned on robot demonstration data and on ordinary internet-scale vision-language tasks, such as visual question answering, at the same time, rather than on robot data alone. The reported effect is generalisation to objects and instructions the robot arm itself never demonstrated — placing an object on a specific printed number, for instance — plus the ability to carry a visible instruction through several reasoning steps before acting, which the authors attribute to knowledge transferred from the web-scale half of the training mixture [1]. That is the developing laboratory’s own account of its own system, a claim rather than an independently audited result — but the mechanism it describes, treating an action as one more token in a shared vocabulary, is architecturally real and has been widely adopted since.

pi-zero takes the second approach: rather than discretizing actions into a token vocabulary, it adds a flow-matching head on top of a pretrained vision-language backbone that outputs continuous action chunks directly, which avoids the resolution loss that comes with discretization and appears to matter for the kind of fine, contact-rich manipulation — folding cloth, assembling a box — where a coarsely quantized action space loses exactly the precision the task needs [2].

Gemini Robotics splits the problem differently again, separating a vision-language-action generalist that directly outputs robot actions from a companion model, Gemini Robotics-ER, whose job is embodied reasoning — spatial and temporal understanding of a scene — rather than motor output. Google DeepMind reports that the resulting system is robust to variation in object type and position, handles environments it was not trained in, follows open-vocabulary instructions, and after fine-tuning can learn new short-horizon tasks from as few as 100 demonstrations and adapt to robot embodiments it was not originally trained on [3]. Again, this is the developer’s own account, a claim rather than an independently reproduced result — but the pattern it exemplifies, splitting reasoning about a scene from acting in it, is a second real design axis alongside the token-versus-continuous choice above.

What all three share is the premise that follows directly from the previous section’s demonstration-scarcity problem: if there is not enough robot-specific data to teach a network what a drawer is from scratch, borrow that knowledge from a much larger corpus that was never collected with robots in mind, and spend the comparatively small robot-specific dataset teaching the network to act rather than to perceive.

Reinforcement learning fine-tunes what imitation cannot

Imitation learning has a ceiling, and the previous article in this series names the structural reason for it: a behavior-cloned policy has only ever seen a demonstrator succeed, so it has no data at all about what to do once it drifts slightly off the demonstrated path, and small errors compound because nothing in training taught the network to recover from them. The mechanism-level answer to that ceiling is not to collect more demonstrations — recoveries are precisely the states a competent demonstrator rarely visits — but to let the robot generate its own corrective experience under a reward signal, which is what reinforcement learning fine-tuning does.

Luo and colleagues built a concrete instance of this pattern for real-world manipulation. Rather than training an RL policy from nothing, their system first uses a small set of teleoperated demonstrations to seed the replay buffer an RL algorithm draws from, trains a binary reward classifier from human-labeled examples of success and failure so the system does not need a hand-engineered reward function, and then runs online reinforcement learning in the real world with a human able to intervene and correct the robot’s behavior live during training, rather than only at the demonstration stage. Across a set of precise and dynamic manipulation tasks, including dual-arm coordination, the resulting policies reached near-perfect success rates within one to two and a half hours of training time — roughly double the success rate and 1.8 times the execution speed of policies trained by imitation learning or by prior reinforcement-learning baselines on the same tasks [11].

Two things about that result matter more than the headline numbers. First, the human’s role changes rather than disappears: a teleoperator who produced every demonstrated trajectory by hand earlier becomes, in the RL stage, an occasional corrector who intervenes only when the policy is already close to a solution — cheaper, more targeted supervision than demonstrating from scratch. Second, the reward classifier does work a hand-written reward function would struggle with on a contact-rich task: judging success from the same camera images the policy itself sees, rather than a hand-coded geometric condition that may not hold for a deformable or occluded object.

This is fine-tuning in the literal sense — it starts from behavior a demonstrator already showed the robot and improves it, rather than discovering a task from an uninformed policy — and it is the piece of the pipeline most directly responsible for pushing performance above what pure imitation reaches on tasks where contact and precision dominate.

Simulation as a rehearsal space

Everything described so far — teleoperated demonstrations, vision-language-action pretraining, real-world reinforcement fine-tuning — is expensive specifically because it happens on hardware, in real time, usually with a human present. Simulation exists in the training loop because it removes both constraints at once: a physics engine can be stepped far faster than real time and run thousands of copies in parallel on a single machine, and nothing breaks when a simulated attempt fails.

Rudin and colleagues demonstrated how large that speed-up can be for locomotion specifically. Training a quadruped, ANYmal, to walk conventionally took substantial wall-clock time; running thousands of simulated instances of the same robot in parallel on a single workstation GPU, alongside a training curriculum built to keep that many parallel agents making progress together, cut the time to train a working flat-terrain locomotion policy to under four minutes, and a policy for uneven terrain to twenty [8]. That is not a modest optimisation; it is roughly the difference between a policy a lab can retrain between coffee breaks and one that requires an overnight run, and it is the concrete mechanism that makes reinforcement learning practical for locomotion at all — RL typically needs orders of magnitude more environment interaction than imitation learning does, which is only affordable if each unit of interaction is nearly free.

Speed alone does not make a simulated policy work on the real robot, because a simulator’s physics is an approximation and the real robot never matches it exactly — heavier here, stiffer there, a motor with slightly more backlash than its model. Domain randomization is the training-time answer to that mismatch: rather than train against one simulated configuration and hope it generalises, vary the simulator’s parameters — visual textures, lighting, object poses, and where applicable friction and mass — widely enough that the real world ends up looking like one more sample from the distribution the policy already learned to handle, rather than a distinct case it has never seen. Tobin and colleagues established the technique’s canonical form for perception: training an object detector entirely on non-realistic simulated images with randomised rendering, using no real images at all, and reaching real-world object localisation accurate to 1.5 centimetres, sufficient to grasp real objects in clutter [7]. The same logic — randomise what the simulator already models — extends to dynamics parameters and is now standard practice across manipulation and locomotion training alike.

A small robot arm on a bench directly below a large wall display rendering a matched simulated version of itself, the physical gripper mid-close on a block while the simulated gripper lags a fraction behind
Figure 3. A policy trained across thousands of randomised simulated variants is only as good as the interface between the render and the room; the transfer bench is where that gap is measured.

Locomotion adds a second use of simulation that manipulation mostly does not need: training the perception side of the loop to fuse sources that disagree. Miki and colleagues trained their terrain-perception encoder — the exteroception-proprioception fusion described earlier in this article — using simulated terrain and simulated sensor noise before deploying the same network on a physical quadruped across natural and urban environments over multiple seasons, including an hour-long Alpine hike completed at the pace recommended for human hikers [9]. The transfer succeeded specifically on the cases domain randomization is built to cover: terrain that visually resembles an obstacle, like tall grass, or terrain a depth sensor can miss entirely, like standing water and fresh snow — deceptions a sufficiently varied simulated training distribution had already prepared the network to discount.

None of this eliminates the sim-to-real gap named elsewhere in this series; it narrows the specific slice of that gap — appearance, and parameters a simulator actually represents — that domain randomization can reach, while leaving contact physics and deformable materials as open problems simulation training does not solve by itself.

From a policy’s target to a joint’s torque

A learned policy, however it was trained, outputs something well short of a torque command: a target end-effector pose, a target joint configuration, a desired body velocity. Turning that target into current through a motor winding is the job of a much older and much less discussed part of the stack — classical control — and it is where the industry’s decades of work on servo systems has not been replaced by learning so much as put underneath it.

The dominant scheme for the last stage is impedance control rather than pure position control, and the distinction matters for exactly the reason a robot has to touch things. A pure position controller tries to force a joint to a commanded angle regardless of what resists it, which is dangerous the instant the arm meets an object, a surface, or a person, because the controller will keep applying more torque to fight the very contact it should be responding to. An impedance controller instead makes the joint behave like a programmable spring and damper around the commanded target:

τ=Kp(qd−q)+Kd(q˙d−q˙)+τff \tau = K_p (q_d - q) + K_d (\dot{q}_d - \dot{q}) + \tau_{ff} ↗

Here qdq_d↗ and qq↗ are the desired and measured joint positions, KpK_p↗ and KdK_d↗ set how stiff and how damped that virtual spring-and-damper is, and τff\tau_{ff}↗ is a feedforward term — gravity compensation, or a reaction-force term handed down from a higher-level controller — added on top. The learned policy’s contribution to this equation is qdq_d↗: a target, updated at whatever rate the policy runs. Everything else in the equation runs at a much higher rate, underneath the policy, and it is what actually decides how the robot responds to unexpected resistance in the interval between two policy decisions. Lower the gains and the joint yields to contact instead of fighting it; that yielding, not any property of the learned policy, is most of what keeps a light unexpected collision from becoming a hard one.

A servo-drive controller board mid-bring-up in a bench vice, one axis's status LED lit steady while a probe lead's crocodile clip rests on a bare test pad beside it
Figure 4. Above the learned policy sits a torque loop most of the training literature never mentions; this board, not the network, is what actually enforces a force limit.

This is also where safety compliance actually lives, and it is worth being specific rather than gestural about it. ISO/TS 15066 governs collaborative robots operating in shared space with people, and one of its modes — power and force limiting — requires the robot to sense contact and hold its interaction force below thresholds set for the body region involved. Ghanbarzadeh and Najafi built a variable impedance controller aimed directly at this requirement, adjusting the gains above in real time so the robot could run at higher speed while keeping worst-case contact force within the standard’s limits, rather than accepting the low fixed speed a constant, conservative impedance would otherwise force onto every motion [12]. The result is a genuine trade rather than a solved problem: stiffer gains track a trajectory more precisely and move faster; softer gains are safer on contact and slower.

The arbitration this article’s framing promises is, at this layer, unglamorous and literal. The learned policy is not asked whether it is allowed to apply unlimited force; it is not consulted at all. Its output is a target the impedance controller is free to miss, on purpose, whenever tracking that target would mean pushing harder against something than the safety-rated gains allow. The override is architectural, not negotiated — which is exactly why it is trusted with a safety case that the learned network, being an unverified statistical function, currently is not.

Whole-body control and the same arbitration at larger scale

Everything in the previous section generalises from a single joint to an entire legged or humanoid robot, but the generalisation is not free: on a manipulator resting on a fixed base, an over-aggressive impedance command in the worst case damages what it touches. On a legged robot, an over-aggressive or simply wrong command can make the whole machine fall over, because standing itself is a continuously re-solved balance problem, not a default state the robot returns to when nothing is commanding it.

Kim and colleagues built a reference architecture for handling that difference, splitting the controller into two layers with different jobs and different rates. A model predictive controller (MPC) plans over a longer horizon using a simplified model of the robot’s dynamics, solving for an optimal profile of ground reaction forces at each foot over the next several planned footsteps. A whole-body controller (WBC) then takes that reaction-force plan and converts it into joint-level torque, position and velocity commands for the robot’s full, much more detailed dynamics, respecting the actual contact constraints — how many feet are on the ground right now, and in what configuration — that the simplified MPC model does not represent in full. On the Mini-Cheetah quadruped, this two-layer split produced dynamic gaits with aerial phases, in which no foot touches the ground at all, and a top speed of 3.7 metres per second across six different gaits and environments [10]. Neither layer alone could do this: the MPC model is too simplified to command individual joints safely, and a whole-body controller with no forward-looking force plan has no way to decide where to place a foot before the robot is already falling toward the answer.

A bipedal test robot suspended in an overhead safety harness on a gantry test stand, one foot just lifting clear of a low platform as a status indicator on its torso lights
Figure 5. A locomotion policy proposes a footstep; a whole-body controller underneath has to keep the machine from falling before the foot ever lands.

This is the same arbitration principle as the impedance equation in the previous section, restated at body scale: a higher-level decision-maker — a classical MPC solving an optimisation problem, or a learned policy trained by imitation or reinforcement learning — proposes a target, and a lower, faster, contact-aware layer realises that target without violating the robot’s physical limits or its balance. The learned locomotion policies described earlier, trained in massively parallel simulation and transferred using domain randomization, typically sit at or above the position the MPC layer occupies in Kim and colleagues’ architecture: the network outputs joint targets or body velocity commands directly, and a lower-level torque or impedance loop is still what converts those targets into current, and still what a safety monitor can override if a foot’s measured load or a joint’s tracking error crosses a threshold the network was never asked to respect.

What changes moving from arms to legged and humanoid platforms is not whether this arbitration exists — it exists in both cases — but how much stability depends on the lower layer getting it right every single cycle: a manipulator that momentarily fails to track a target drops what it was holding; a legged robot that fails to track a balance-critical target falls.

What to take away

Trace the loop from end to end and the architecture is more legible than the individual papers make it look. Sensing and state estimation supply a fused, current estimate of the world. A policy — pretrained on internet-scale vision-language data, fitted to teleoperated demonstrations by imitation learning, and in a growing number of systems further sharpened by real-world reinforcement fine-tuning — turns that estimate into a target, running as fast as its own inference cost allows. Simulation, sped up by orders of magnitude through massive parallelism and widened by domain randomization, is where as much of that policy as possible is trained and tested before any of it touches a real actuator. And beneath all of it, a much older layer of classical control — impedance control on a single joint, whole-body control and model predictive control across an entire legged body — converts the policy’s target into torque while retaining, deliberately and architecturally, the authority to refuse it.

That last point is the one worth carrying forward. The learned parts of this stack get faster, more general, and better at handling instructions and objects they were never explicitly shown. The part of the stack that is allowed to override them, when contact, force or balance demand it, has stayed classical, verifiable and comparatively unglamorous — and every system described in this article was built to keep it that way.