← All parts of this equation

Equation 5 · Part 14 · Comparing the Main Approaches to Robotics and Embodied AI

Ending index or upper bound: H

a1:H∗=arg⁡max⁡a1:H  Ep^θ[∑t=1Hγt r(st,at)]a^{*}_{1:H} = \arg\max_{a_{1:H}} \; \mathbb{E}_{\hat{p}_\theta} \left[ \sum_{t=1}^{H} \gamma^{t} \, r(s_t, a_t) \right]
HH

What this part means

This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.

Its job in the formula

H appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

The passage around this formula

…queried and its predictions checked against reality independent of the policy it supports; the cost is that policy quality is now bounded by model accuracy, and errors in p^θ\hat{p}_\theta compound over the planning horizon H in ways that are hard to detect from the outside [ 6 ] . Classical robotics has practiced a version of this for decades without calling it “model-based reinforcement learning”: model-predictive control on the MIT Cheetah 3 solves a convex optimization over ground reaction forces against a…

Read this part in the article →

Learn the underlying idea

Σ adds a collection of terms. Π multiplies them. The lower and upper labels tell you which terms belong to the collection.

Open the illustrated sums and products: repeat an operation over an index guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.