Equation 1 · Part 13 · The Hardest Unsolved Problems in AI Agent Architecture
superscript
superscript
What this part means
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
Its job in the formula
A raised mark can be a power or an index. Its position and the surrounding notation determine which.
Full expression→superscript→Article meaning
The passage around this formula
Reinforcement learning’s standard tool for turning a sequence of actions and a single delayed reward into a training signal is the policy gradient, and its textbook reward-to-go form is . Read the inner sum literally: every action in the trajectory is credited with everything that happens from that point on, not with its own specific contribution. When T is small and rewards are dense, that crude attribution washes out quickly. When T is large and the reward is a single terminal signal — a multi-step coding task that either compiles and passes its tests or does not, a multi-turn support conversation that either resolves the case or does not — the same sum assigns identical…
Learn the underlying idea
An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.
Open the illustrated exponents: repeated multiplication and powers guide →
Sources cited in the article section
- [2] ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL ↗
- [3] Reflexion: Language Agents with Verbal Reinforcement Learning ↗
These citations provide research context; check each source for the exact claim it supports.