← All parts of this equation

Equation 5 · Part 3 · Comparing the Main Approaches to Robotics and Embodied AI

Symbol E_hatp_θ

a1:H∗=arg⁡max⁡a1:H  Ep^θ[∑t=1Hγt r(st,at)]a^{*}_{1:H} = \arg\max_{a_{1:H}} \; \mathbb{E}_{\hat{p}_\theta} \left[ \sum_{t=1}^{H} \gamma^{t} \, r(s_t, a_t) \right]
Ep^θ\mathbb{E}_{\hat{p}_\theta}

What this part means

EhE_hatp_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Its job in the formula

EhE_hatp_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

The passage around this formula

Nothing in this update requires knowing or estimating p(s' ∣\mid s, a) ; it only requires having experienced (s, a, r, s') . A model-based approach instead fits an explicit dynamics model p^θ(s′∣s,a)\hat{p}_\theta(s' \mid s, a) — a “world model” — and plans against it directly: a1:H∗=arg⁡max⁡a1:H  Ep^θ[∑t=1Hγt r(st,at)]a^{*}_{1:H} = \arg\max_{a_{1:H}} \; \mathbb{E}_{\hat{p}_\theta} \left[ \sum_{t=1}^{H} \gamma^{t} \, r(s_t, a_t) \right]. A comprehensive survey of the model-based literature frames the trade this way: fitting and planning against p^θ\hat{p}_\theta typically buys sample efficiency, because every transition teaches the model something reusable across many hypothetical future plans rather than updating one value estimate, and it buys interpretability, because the model can be queried and its predictions checked against reality…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.