Equation 3 · Part 13 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness
superscript
superscript
What this part means
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
Its job in the formula
A raised mark can be a power or an index. Its position and the surrounding notation determine which.
Full expression→superscript→Article meaning
The passage around this formula
Policy optimisation. The fine-tuned model is then optimised, typically with PPO [ 6 ] , to maximise the reward model’s score, with a penalty on divergence from the starting policy. InstructGPT’s published objective adds a per-token KL penalty from the supervised model and, in the PPO-ptx variant, a term mixing in pretraining gradients: . where sets the strength of the KL penalty and the pretraining mixture, with set to zero for the plain PPO models [ 4 ] . Direct preference optimisation later showed that the explicit reward model can be dispensed with entirely — the optimal policy under this objective has a closed form, so the same problem can be solved…
Learn the underlying idea
An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.
Open the illustrated exponents: repeated multiplication and powers guide →
Sources cited in the surrounding passage
- [6] Proximal Policy Optimization Algorithms ↗
- [4] Training Language Models to Follow Instructions with Human Feedback ↗
- [15] Direct Preference Optimization: Your Language Model is Secretly a Reward Model ↗
These citations provide research context; check each source for the exact claim it supports.