Equation 23 · Part 11 · The Main Technical Approaches to AI Alignment, Compared
Denominator: std(r_1,ldots,r_G)
What this part means
The complete quantity below the fraction bar; it must be nonzero for this division.
Its job in the formula
std(,ldots,) occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Full expression→Denominator: std(r_1,ldots,r_G)→Article meaning
The passage around this formula
RLVR is the newest of the six and the only one that does not fit a learned model of anyone’s judgment at all. Where every technique above trains a proxy — a reward model, an AI judge, a decomposition scheme, a weak label — RLVR restricts itself to tasks where the reward can be computed directly and automatically: a unit test passes, a final numeric answer matches, a proof checker accepts. Shao and colleagues’ DeepSeekMath paper introduced Group Relative Policy Optimization, the reinforcement learning algorithm used throughout most subsequent RLVR work, replacing the learned value function used in standard policy-gradient methods with a group-relative advantage estimated directly from a batch…
Learn the underlying idea
A fraction a/b means a divided by b. The top number is the numerator; the bottom number is the denominator, and it cannot be zero.
Open the illustrated fractions: division written vertically guide →
Sources cited in the surrounding passage
- [13] DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models ↗
- [15] Tulu 3: Pushing Frontiers in Open Language Model Post-Training ↗
- [14] DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning ↗
These citations provide research context; check each source for the exact claim it supports.