← Mathematical compendium

Published equation contexts

∇θJ(θ)=Eπ ⁣[∑t=0T∇θlog⁡πθ(at∣st)(∑t′≥trt′)]\nabla_\theta J(\theta) = \mathbb{E}_\pi\!\left[\sum_{t=0}^{T} \nabla_\theta \log \pi_\theta(a_t \mid s_t) \left(\sum_{t' \ge t} r_{t'}\right)\right]

Why this formula appears here

Reinforcement learning’s standard tool for turning a sequence of actions and a single delayed reward into a training signal is the policy gradient, and its textbook reward-to-go form is ∇θJ(θ)=Eπ ⁣[∑t=0T∇θlog⁡πθ(at∣st)(∑t′≥trt′)]\nabla_\theta J(\theta) = \mathbb{E}_\pi\!\left[\sum_{t=0}^{T} \nabla_\theta \log \pi_\theta(a_t \mid s_t) \left(\sum_{t' \ge t} r_{t'}\right)\right]. Read the inner sum literally: every action in the trajectory is credited with everything that happens from that point on, not with its own specific contribution. When T is small and rewards are dense, that crude attribution washes out quickly. When T is large and the reward is a single terminal signal — a multi-step coding task that either compiles and passes its tests or does not, a multi-turn support conversation that either resolves the case or does not — the same sum assigns identical…

Read the full article-specific guide →

Read the representative guide

Eπ\mathbb{E}_\pi

Symbol E_pi

EpE_pi appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
tt

Symbol t

t appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Read this term in its guide →
πθ\pi_\theta

Symbol pi_θ

pi_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
ata_t

Symbol a_t

ata_t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
sts_t

Symbol s_t

sts_t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
rt′r_{t'}

Symbol r_t'

rtr_t' appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
t=0t=0

Starting index or lower bound: t=0

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Read this term in its guide →
TT

Ending index or upper bound: T

This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.

Read this term in its guide →
t′≥tt' \ge t

Starting index or lower bound: t' ≥ t

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

∇θJ(θ)=Eπ ⁣[∑t=0T∇θlog⁡πθ(at∣st)(∑t′≥trt′)].\nabla_\theta J(\theta) = \mathbb{E}_\pi\!\left[\sum_{t=0}^{T} \nabla_\theta \log \pi_\theta(a_t \mid s_t) \left(\sum_{t' \ge t} r_{t'}\right)\right].

Equation 1 · AI Agents & Systems

The Hardest Unsolved Problems in AI Agent Architecture

This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions.

Reinforcement learning’s standard tool for turning a sequence of actions and a single delayed reward into a training signal is the policy gradient, and its textbook reward-to-go form is ∇θJ(θ)=Eπ ⁣[∑t=0T∇θlog⁡πθ(at∣st)(∑t′≥trt′)]\nabla_\theta J(\theta) = \mathbb{E}_\pi\!\left[\sum_{t=0}^{T} \nabla_\theta \log \pi_\theta(a_t \mid s_t) \left(\sum_{t' \ge t} r_{t'}\right)\right]. Read the inner sum literally: every action in the trajectory is credited with everything that happens from that point on, not with its own specific contribution. When T is small and rewards are dense, that crude attribution washes out quickly. When T is large and the reward is a single terminal signal — a multi-step coding task that either compiles and passes its tests or does not, a multi-turn support conversation that either resolves the case or does not — the same sum assigns identical…

Meanings in this article

Equation guide → · Article →