Equation 1 · Part 14 · The Hardest Unsolved Problems in AI Agent Architecture
Starting index or lower bound: t=0
What this part means
This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
Its job in the formula
t=0 appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Full expression→Starting index or lower bound: t=0→Article meaning
The passage around this formula
Reinforcement learning’s standard tool for turning a sequence of actions and a single delayed reward into a training signal is the policy gradient, and its textbook reward-to-go form is . Read the inner sum literally: every action in the trajectory is credited with everything that happens from that point on, not with its own specific contribution. When T is small and rewards are dense, that crude attribution washes out quickly. When T is large and the reward is a single terminal signal — a multi-step coding task that either compiles and passes its tests or does not, a multi-turn support conversation that either resolves the case or does not — the same sum assigns identical…
Learn the underlying idea
Σ adds a collection of terms. Π multiplies them. The lower and upper labels tell you which terms belong to the collection.
Open the illustrated sums and products: repeat an operation over an index guide →
Sources cited in the article section
- [2] ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL ↗
- [3] Reflexion: Language Agents with Verbal Reinforcement Learning ↗
These citations provide research context; check each source for the exact claim it supports.