Equation 9 · Part 1 · From n-Grams to Reasoning Models: A Technical History of the Language Model
Symbol r
What this part means
r is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
r is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol r→Article meaning
The passage around this formula
where r is not a learned preference model but a program: a unit test that passes, a numerical answer that matches, a proof that checks. Because r is exact, it cannot be gamed the way a learned reward model can, though it is only available where correctness is mechanically decidable.
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the article section
- [20] Chain-of-Thought Prompting Elicits Reasoning in Large Language Models ↗
- [21] Training Verifiers to Solve Math Word Problems ↗
- [22] Tülu 3: Pushing Frontiers in Open Language Model Post-Training ↗
- [23] DeepSeek-R1 Incentivizes Reasoning in LLMs through Reinforcement Learning ↗
These citations provide research context; check each source for the exact claim it supports.