← All parts of this equation

Equation 1 · Part 9 · The Hardest Unsolved Problems in AI Agent Architecture

Symbol r_t'

∇θJ(θ)=Eπ ⁣[∑t=0T∇θlog⁡πθ(at∣st)(∑t′≥trt′)].\nabla_\theta J(\theta) = \mathbb{E}_\pi\!\left[\sum_{t=0}^{T} \nabla_\theta \log \pi_\theta(a_t \mid s_t) \left(\sum_{t' \ge t} r_{t'}\right)\right].
rt′r_{t'}

What this part means

rtr_t' appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Its job in the formula

rtr_t' appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

The passage around this formula

Reinforcement learning’s standard tool for turning a sequence of actions and a single delayed reward into a training signal is the policy gradient, and its textbook reward-to-go form is ∇θJ(θ)=Eπ ⁣[∑t=0T∇θlog⁡πθ(at∣st)(∑t′≥trt′)]\nabla_\theta J(\theta) = \mathbb{E}_\pi\!\left[\sum_{t=0}^{T} \nabla_\theta \log \pi_\theta(a_t \mid s_t) \left(\sum_{t' \ge t} r_{t'}\right)\right]. Read the inner sum literally: every action in the trajectory is credited with everything that happens from that point on, not with its own specific contribution. When T is small and rewards are dense, that crude attribution washes out quickly. When T is large and the reward is a single terminal signal — a multi-step coding task that either compiles and passes its tests or does not, a multi-turn support conversation that either resolves the case or does not — the same sum assigns identical…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.