← All parts of this equation

Equation 5 · Part 1 · Comparing the Main Approaches to Robotics and Embodied AI

Symbol a^*_1:H

a1:H∗=arg⁡max⁡a1:H  Ep^θ[∑t=1Hγt r(st,at)]a^{*}_{1:H} = \arg\max_{a_{1:H}} \; \mathbb{E}_{\hat{p}_\theta} \left[ \sum_{t=1}^{H} \gamma^{t} \, r(s_t, a_t) \right]
a1:H∗a^{*}_{1:H}

What this part means

a1∗a^*_1:H is computed from the expected values combined on the right.

Its job in the formula

a1∗a^*_1:H is computed from the expected values combined on the right.

The passage around this formula

Nothing in this update requires knowing or estimating p(s' ∣\mid s, a) ; it only requires having experienced (s, a, r, s') . A model-based approach instead fits an explicit dynamics model p^θ(s′∣s,a)\hat{p}_\theta(s' \mid s, a) — a “world model” — and plans against it directly: a1:H∗=arg⁡max⁡a1:H  Ep^θ[∑t=1Hγt r(st,at)]a^{*}_{1:H} = \arg\max_{a_{1:H}} \; \mathbb{E}_{\hat{p}_\theta} \left[ \sum_{t=1}^{H} \gamma^{t} \, r(s_t, a_t) \right]. A comprehensive survey of the model-based literature frames the trade this way: fitting and planning against p^θ\hat{p}_\theta typically buys sample efficiency, because every transition teaches the model something reusable across many hypothetical future plans rather than updating one value estimate, and it buys interpretability, because the model can be queried and its predictions checked against reality…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.