Symbol hat A_i,t
hat ,t is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
RLVR is the newest of the six and the only one that does not fit a learned model of anyone’s judgment at all. Where every technique above trains a proxy — a reward model, an AI judge, a decomposition scheme, a weak label — RLVR restricts itself to tasks where the reward can be computed directly and automatically: a unit test passes, a final numeric answer matches, a proof checker accepts. Shao and colleagues’ DeepSeekMath paper introduced Group Relative Policy Optimization, the reinforcement learning algorithm used throughout most subsequent RLVR work, replacing the learned value function used in standard policy-gradient methods with a group-relative advantage estimated directly from a batch…
hat ,t is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →tilde is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Read this term in its guide →is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →The complete quantity above the fraction bar.
Read this term in its guide →The complete quantity below the fraction bar; it must be nonzero for this division.
Read this term in its guide →With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 23 · AI Safety
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
RLVR is the newest of the six and the only one that does not fit a learned model of anyone’s judgment at all. Where every technique above trains a proxy — a reward model, an AI judge, a decomposition scheme, a weak label — RLVR restricts itself to tasks where the reward can be computed directly and automatically: a unit test passes, a final numeric answer matches, a proof checker accepts. Shao and colleagues’ DeepSeekMath paper introduced Group Relative Policy Optimization, the reinforcement learning algorithm used throughout most subsequent RLVR work, replacing the learned value function used in standard policy-gradient methods with a group-relative advantage estimated directly from a batch…
Equation guide → · Article →