← Mathematical compendium

Published equation contexts

A^i,t  =  r~i  =  ri−mean({r1,…,rG})std({r1,…,rG})\hat A_{i,t} \;=\; \tilde r_i \;=\; \frac{r_i - \mathrm{mean}(\{r_1,\ldots,r_G\})}{\mathrm{std}(\{r_1,\ldots,r_G\})}

Why this formula appears here

RLVR is the newest of the six and the only one that does not fit a learned model of anyone’s judgment at all. Where every technique above trains a proxy — a reward model, an AI judge, a decomposition scheme, a weak label — RLVR restricts itself to tasks where the reward can be computed directly and automatically: a unit test passes, a final numeric answer matches, a proof checker accepts. Shao and colleagues’ DeepSeekMath paper introduced Group Relative Policy Optimization, the reinforcement learning algorithm used throughout most subsequent RLVR work, replacing the learned value function used in standard policy-gradient methods with a group-relative advantage estimated directly from a batch…

Read the full article-specific guide →

Read the representative guide

A^i,t\hat A_{i,t}

Symbol hat A_i,t

hat AiA_i,t is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →
r~i\tilde r_i

Symbol tilde r_i

tilde rir_i is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →
ri−mean({r1,…,rG})r_i - \mathrm{mean}(\{r_1,\ldots,r_G\})

Numerator: r_i - mean(r_1,ldots,r_G)

The complete quantity above the fraction bar.

Read this term in its guide →
std({r1,…,rG})\mathrm{std}(\{r_1,\ldots,r_G\})

Denominator: std(r_1,ldots,r_G)

The complete quantity below the fraction bar; it must be nonzero for this division.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

A^i,t  =  r~i  =  ri−mean({r1,…,rG})std({r1,…,rG}),\hat A_{i,t} \;=\; \tilde r_i \;=\; \frac{r_i - \mathrm{mean}(\{r_1,\ldots,r_G\})}{\mathrm{std}(\{r_1,\ldots,r_G\})},

Equation 23 · AI Safety

The Main Technical Approaches to AI Alignment, Compared

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

RLVR is the newest of the six and the only one that does not fit a learned model of anyone’s judgment at all. Where every technique above trains a proxy — a reward model, an AI judge, a decomposition scheme, a weak label — RLVR restricts itself to tasks where the reward can be computed directly and automatically: a unit test passes, a final numeric answer matches, a proof checker accepts. Shao and colleagues’ DeepSeekMath paper introduced Group Relative Policy Optimization, the reinforcement learning algorithm used throughout most subsequent RLVR work, replacing the learned value function used in standard policy-gradient methods with a group-relative advantage estimated directly from a batch…

Equation guide → · Article →