← Mathematical compendium

Published equation contexts

s^t+1=fθ(st,at),o^t+1∼pθ(o∣s^t+1)\hat{s}_{t+1} = f_\theta(s_t, a_t), \qquad \hat{o}_{t+1} \sim p_\theta(o \mid \hat{s}_{t+1})

Why this formula appears here

Dreamer-style world models learn a compact latent-dynamics model — a recurrent state-space model that predicts both a stochastic and a deterministic component of future latent state — and then train a policy by backpropagating through imagined rollouts inside that latent space, rather than through real or simulated environment steps directly [ 8 ] . DayDreamer took that specific mechanism, previously demonstrated mostly in video-game and simulated benchmarks, and applied it to physical robots learning online, without a simulator and without human demonstrations, reporting that four different physical robots (a quadruped and several manipulator arms) learned locomotion and manipulation skills…

Read the full article-specific guide →

Read the representative guide

s^t+1\hat{s}_{t+1}

Symbol hats_t+1

hatsts_t+1 is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →
fθf_\theta

Symbol f_θ

learned from a finite amount of real interaction and is wrong in ways that compound over an imagined rollout’s horizon — which is precisely the boundary the next section is about.

Read this term in its guide →
o^t+1\hat{o}_{t+1}

Symbol hato_t+1

hatoto_t+1 is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

s^t+1=fθ(st,at),o^t+1∼pθ(o∣s^t+1)\hat{s}_{t+1} = f_\theta(s_t, a_t), \qquad \hat{o}_{t+1} \sim p_\theta(o \mid \hat{s}_{t+1})

Equation 1 · Robotics

How Robotics and Embodied AI Actually Work

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Dreamer-style world models learn a compact latent-dynamics model — a recurrent state-space model that predicts both a stochastic and a deterministic component of future latent state — and then train a policy by backpropagating through imagined rollouts inside that latent space, rather than through real or simulated environment steps directly [ 8 ] . DayDreamer took that specific mechanism, previously demonstrated mostly in video-game and simulated benchmarks, and applied it to physical robots learning online, without a simulator and without human demonstrations, reporting that four different physical robots (a quadruped and several manipulator arms) learned locomotion and manipulation skills…

Meanings in this article

  • fθf_\theta: learned from a finite amount of real interaction and is wrong in ways that compound over an imagined rollout’s horizon — which is precisely the boundary the next section is about.
Equation guide → · Article →