Equation 23 · Part 8 · The Main Technical Approaches to AI Alignment, Compared
subtraction
subtraction
What this part means
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
Its job in the formula
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
Full expression→subtraction→Article meaning
The passage around this formula
RLVR is the newest of the six and the only one that does not fit a learned model of anyone’s judgment at all. Where every technique above trains a proxy — a reward model, an AI judge, a decomposition scheme, a weak label — RLVR restricts itself to tasks where the reward can be computed directly and automatically: a unit test passes, a final numeric answer matches, a proof checker accepts. Shao and colleagues’ DeepSeekMath paper introduced Group Relative Policy Optimization, the reinforcement learning algorithm used throughout most subsequent RLVR work, replacing the learned value function used in standard policy-gradient methods with a group-relative advantage estimated directly from a batch…
Learn the underlying idea
Addition combines quantities; subtraction measures the signed difference between them. Parentheses show what is combined before the rest of the expression is evaluated.
Open the illustrated addition and subtraction in an equation guide →
Sources cited in the surrounding passage
- [13] DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models ↗
- [15] Tulu 3: Pushing Frontiers in Open Language Model Post-Training ↗
- [14] DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning ↗
These citations provide research context; check each source for the exact claim it supports.