Equation 3 · The Hardest Unsolved Problems in AI Agent Architecture
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the when. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Read the inner sum literally: every action in the trajectory is credited with everything that happens from that point on, not with its own specific contribution. When T is small and rewards are dense, that crude attribution washes out quickly. When T is large and the reward is a single terminal signal — a multi-step coding task that either compiles and passes its tests or does not, a multi-turn support conversation that either resolves the case or does not — the same sum assigns identical credit to a decisive action, a neutral one, and a mistake that happened to be recovered from later. The variance of the resulting gradient estimate grows with the length of the trajectory, and no…
Read the full surrounding passage
Read the inner sum literally: every action in the trajectory is credited with everything that happens from that point on, not with its own specific contribution. When T is small and rewards are dense, that crude attribution washes out quickly. When T is large and the reward is a single terminal signal — a multi-step coding task that either compiles and passes its tests or does not, a multi-turn support conversation that either resolves the case or does not — the same sum assigns identical credit to a decisive action, a neutral one, and a mistake that happened to be recovered from later. The variance of the resulting gradient estimate grows with the length of the trajectory, and no architectural trick changes that arithmetic; it only changes how the trajectory is chunked before credit is assigned.
Sources cited in the article section
- [2] ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL ↗
- [3] Reflexion: Language Agents with Verbal Reinforcement Learning ↗
These citations give research context. Read each source to check which claims it supports.
Return to The Hardest Unsolved Problems in AI Agent Architecture