← All parts of this equation

Equation 3 · Part 10 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness

addition

E[rθ(x,y)]−β E[DKL ⁣(πRL ∥ πSFT)]+γ E[log⁡πRL(x)],\mathbb{E}\big[r_\theta(x,y)\big] - \beta\,\mathbb{E}\big[D_{\mathrm{KL}}\!\left(\pi^{\mathrm{RL}} \,\|\, \pi^{\mathrm{SFT}}\right)\big] + \gamma\,\mathbb{E}\big[\log \pi^{\mathrm{RL}}(x)\big],
addition

What this part means

Add the term after the plus sign to the term or group before it.

Its job in the formula

Add the term after the plus sign to the term or group before it.

The passage around this formula

Policy optimisation. The fine-tuned model is then optimised, typically with PPO [ 6 ] , to maximise the reward model’s score, with a penalty on divergence from the starting policy. InstructGPT’s published objective adds a per-token KL penalty from the supervised model and, in the PPO-ptx variant, a term mixing in pretraining gradients: E[rθ(x,y)]−β E[DKL ⁣(πRL ∥ πSFT)]+γ E[log⁡πRL(x)]\mathbb{E}\big[r_\theta(x,y)\big] - \beta\,\mathbb{E}\big[D_{\mathrm{KL}}\!\left(\pi^{\mathrm{RL}} \,\|\, \pi^{\mathrm{SFT}}\right)\big] + \gamma\,\mathbb{E}\big[\log \pi^{\mathrm{RL}}(x)\big]. where β\beta sets the strength of the KL penalty and γ\gamma the pretraining mixture, with γ\gamma set to zero for the plain PPO models [ 4 ] . Direct preference optimisation later showed that the explicit reward model can be dispensed with entirely — the optimal policy under this objective has a closed form, so the same problem can be solved…

Read this part in the article →

Learn the underlying idea

Addition combines quantities; subtraction measures the signed difference between them. Parentheses show what is combined before the rest of the expression is evaluated.

Open the illustrated addition and subtraction in an equation guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.