Equation 5 · Part 1 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness
Symbol gamma
What this part means
the pretraining mixture.
Its job in the formula
gamma is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol gamma→Article meaning
Where the article explains it
where sets the strength of the KL penalty and the pretraining mixture, with set to zero for the plain PPO models [ 4 ] .
The passage around this formula
where sets the strength of the KL penalty and the pretraining mixture, with set to zero for the plain PPO models [ 4 ] . Direct preference optimisation later showed that the explicit reward model can be dispensed with entirely — the optimal policy under this objective has a closed form, so the same problem can be solved with a classification loss on the preference pairs [ 15 ] . That is an important simplification, and it changes nothing about the present argument. DPO removes the reward network; it does not remove the reward. The objective is still rater approval, and the KL term is still there, folded into the loss.
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the surrounding passage
- [4] Training Language Models to Follow Instructions with Human Feedback ↗
- [15] Direct Preference Optimization: Your Language Model is Secretly a Reward Model ↗
These citations provide research context; check each source for the exact claim it supports.