Published equation contexts
Why this formula appears here
Reinforcement learning’s standard tool for turning a sequence of actions and a single delayed reward into a training signal is the policy gradient, and its textbook reward-to-go form is . Read the inner sum literally: every action in the trajectory is credited with everything that happens from that point on, not with its own specific contribution. When T is small and rewards are dense, that crude attribution washes out quickly. When T is large and the reward is a single terminal signal — a multi-step coding task that either compiles and passes its tests or does not, a multi-turn support conversation that either resolves the case or does not — the same sum assigns identical…
Read the representative guide
Symbol E_pi
i appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →Symbol t
t appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Read this term in its guide →Symbol pi_θ
pi_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →Symbol a_t
appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →Symbol s_t
appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →Symbol r_t'
' appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →Starting index or lower bound: t=0
This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
Read this term in its guide →Ending index or upper bound: T
This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.
Read this term in its guide →Starting index or lower bound: t' ≥ t
This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
Read this term in its guide →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
Research cited beside this formula
Published contexts (1)
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 1 · AI Agents & Systems
The Hardest Unsolved Problems in AI Agent Architecture
This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions.
Reinforcement learning’s standard tool for turning a sequence of actions and a single delayed reward into a training signal is the policy gradient, and its textbook reward-to-go form is . Read the inner sum literally: every action in the trajectory is credited with everything that happens from that point on, not with its own specific contribution. When T is small and rewards are dense, that crude attribution washes out quickly. When T is large and the reward is a single terminal signal — a multi-step coding task that either compiles and passes its tests or does not, a multi-turn support conversation that either resolves the case or does not — the same sum assigns identical…