Equation 4 · Part 1 · Comparing the Main Approaches to Robotics and Embodied AI
Symbol hatp_θ
What this part means
hatp_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
hatp_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol hatp_θ→Article meaning
The passage around this formula
Nothing in this update requires knowing or estimating p(s' s, a) ; it only requires having experienced (s, a, r, s') . A model-based approach instead fits an explicit dynamics model — a “world model” — and plans against it directly:
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the article section
- [6] Model-Based Reinforcement Learning: A Survey ↗
- [10] Dynamic Locomotion in the MIT Cheetah 3 Through Convex Model-Predictive Control ↗
- [9] Mastering Diverse Domains through World Models ↗
- [3] Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning ↗
These citations provide research context; check each source for the exact claim it supports.