Equation 3 · Part 10 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness
addition
addition
What this part means
Add the term after the plus sign to the term or group before it.
Its job in the formula
Add the term after the plus sign to the term or group before it.
Full expression→addition→Article meaning
The passage around this formula
Policy optimisation. The fine-tuned model is then optimised, typically with PPO [ 6 ] , to maximise the reward model’s score, with a penalty on divergence from the starting policy. InstructGPT’s published objective adds a per-token KL penalty from the supervised model and, in the PPO-ptx variant, a term mixing in pretraining gradients: . where sets the strength of the KL penalty and the pretraining mixture, with set to zero for the plain PPO models [ 4 ] . Direct preference optimisation later showed that the explicit reward model can be dispensed with entirely — the optimal policy under this objective has a closed form, so the same problem can be solved…
Learn the underlying idea
Addition combines quantities; subtraction measures the signed difference between them. Parentheses show what is combined before the rest of the expression is evaluated.
Open the illustrated addition and subtraction in an equation guide →
Sources cited in the surrounding passage
- [6] Proximal Policy Optimization Algorithms ↗
- [4] Training Language Models to Follow Instructions with Human Feedback ↗
- [15] Direct Preference Optimization: Your Language Model is Secretly a Reward Model ↗
These citations provide research context; check each source for the exact claim it supports.