Equation 24 · Part 1 · The Main Technical Approaches to AI Alignment, Compared
Symbol G
What this part means
G is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
G is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol G→Article meaning
The passage around this formula
normalizing each sampled response’s reward against the mean and standard deviation of a group of G responses to the same prompt rather than against a separately trained critic network [ 13 ] . Lambert and colleagues, building the fully open Tulu 3 post-training recipe, named the general approach explicitly, describing it as “a novel method we call Reinforcement Learning with Verifiable…
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the surrounding passage
- [13] DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models ↗
- [15] Tulu 3: Pushing Frontiers in Open Language Model Post-Training ↗
- [14] DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning ↗
These citations provide research context; check each source for the exact claim it supports.