← All parts of this equation

Equation 5 · Part 6 · Comparing the Main Approaches to Robotics and Embodied AI

Symbol gamma^t

a1:H∗=arg⁡max⁡a1:H  Ep^θ[∑t=1Hγt r(st,at)]a^{*}_{1:H} = \arg\max_{a_{1:H}} \; \mathbb{E}_{\hat{p}_\theta} \left[ \sum_{t=1}^{H} \gamma^{t} \, r(s_t, a_t) \right]
γt\gamma^{t}

What this part means

gammata^t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Its job in the formula

gammata^t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

The passage around this formula

Nothing in this update requires knowing or estimating p(s' ∣\mid s, a) ; it only requires having experienced (s, a, r, s') . A model-based approach instead fits an explicit dynamics model p^θ(s′∣s,a)\hat{p}_\theta(s' \mid s, a) — a “world model” — and plans against it directly: a1:H∗=arg⁡max⁡a1:H  Ep^θ[∑t=1Hγt r(st,at)]a^{*}_{1:H} = \arg\max_{a_{1:H}} \; \mathbb{E}_{\hat{p}_\theta} \left[ \sum_{t=1}^{H} \gamma^{t} \, r(s_t, a_t) \right]. A comprehensive survey of the model-based literature frames the trade this way: fitting and planning against p^θ\hat{p}_\theta typically buys sample efficiency, because every transition teaches the model something reusable across many hypothetical future plans rather than updating one value estimate, and it buys interpretability, because the model can be queried and its predictions checked against reality…

Read this part in the article →

Learn the underlying idea

An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.

Open the illustrated exponents: repeated multiplication and powers guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.