Symbol hats_t+1
hat+1 is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
Dreamer-style world models learn a compact latent-dynamics model — a recurrent state-space model that predicts both a stochastic and a deterministic component of future latent state — and then train a policy by backpropagating through imagined rollouts inside that latent space, rather than through real or simulated environment steps directly [ 8 ] . DayDreamer took that specific mechanism, previously demonstrated mostly in video-game and simulated benchmarks, and applied it to physical robots learning online, without a simulator and without human demonstrations, reporting that four different physical robots (a quadruped and several manipulator arms) learned locomotion and manipulation skills…
hat+1 is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →learned from a finite amount of real interaction and is wrong in ways that compound over an imagined rollout’s horizon — which is precisely the boundary the next section is about.
Read this term in its guide →is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →hat+1 is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →p_θ is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →o is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 1 · Robotics
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
Dreamer-style world models learn a compact latent-dynamics model — a recurrent state-space model that predicts both a stochastic and a deterministic component of future latent state — and then train a policy by backpropagating through imagined rollouts inside that latent space, rather than through real or simulated environment steps directly [ 8 ] . DayDreamer took that specific mechanism, previously demonstrated mostly in video-game and simulated benchmarks, and applied it to physical robots learning online, without a simulator and without human demonstrations, reporting that four different physical robots (a quadruped and several manipulator arms) learned locomotion and manipulation skills…