← All parts of this equation

Equation 23 · Part 10 · The Main Technical Approaches to AI Alignment, Compared

Numerator: r_i - mean(r_1,ldots,r_G)

A^i,t  =  r~i  =  ri−mean({r1,…,rG})std({r1,…,rG}),\hat A_{i,t} \;=\; \tilde r_i \;=\; \frac{r_i - \mathrm{mean}(\{r_1,\ldots,r_G\})}{\mathrm{std}(\{r_1,\ldots,r_G\})},
ri−mean({r1,…,rG})r_i - \mathrm{mean}(\{r_1,\ldots,r_G\})

What this part means

The complete quantity above the fraction bar.

Its job in the formula

rir_i - mean(r1r_1,ldots,rGr_G) occurs above the fraction bar. The numerator is divided by the entire denominator below it.

The passage around this formula

RLVR is the newest of the six and the only one that does not fit a learned model of anyone’s judgment at all. Where every technique above trains a proxy — a reward model, an AI judge, a decomposition scheme, a weak label — RLVR restricts itself to tasks where the reward can be computed directly and automatically: a unit test passes, a final numeric answer matches, a proof checker accepts. Shao and colleagues’ DeepSeekMath paper introduced Group Relative Policy Optimization, the reinforcement learning algorithm used throughout most subsequent RLVR work, replacing the learned value function used in standard policy-gradient methods with a group-relative advantage estimated directly from a batch…

Read this part in the article →

Learn the underlying idea

A fraction a/b means a divided by b. The top number is the numerator; the bottom number is the denominator, and it cannot be zero.

Open the illustrated fractions: division written vertically guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.