Equation 1 · Part 15 · The Hardest Unsolved Problems in AI Agent Architecture
Ending index or upper bound: T
What this part means
This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.
Its job in the formula
T appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Full expression→Ending index or upper bound: T→Article meaning
The passage around this formula
…textbook reward-to-go form is . Read the inner sum literally: every action in the trajectory is credited with everything that happens from that point on, not with its own specific contribution. When T is small and rewards are dense, that crude attribution washes out quickly. When T is large and the reward is a single terminal signal — a multi-step coding task that either compiles and passes its tests or does not, a multi-turn support conversation that either resolves the case or does not — the…
Learn the underlying idea
Σ adds a collection of terms. Π multiplies them. The lower and upper labels tell you which terms belong to the collection.
Open the illustrated sums and products: repeat an operation over an index guide →
See this notation across published equations →
Sources cited in the article section
- [2] ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL ↗
- [3] Reflexion: Language Agents with Verbal Reinforcement Learning ↗
These citations provide research context; check each source for the exact claim it supports.