← Back to article

Equation 3 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness

What does this equation mean?

E[rθ(x,y)]−β E[DKL ⁣(πRL ∥ πSFT)]+γ E[log⁡πRL(x)],\mathbb{E}\big[r_\theta(x,y)\big] - \beta\,\mathbb{E}\big[D_{\mathrm{KL}}\!\left(\pi^{\mathrm{RL}} \,\|\, \pi^{\mathrm{SFT}}\right)\big] + \gamma\,\mathbb{E}\big[\log \pi^{\mathrm{RL}}(x)\big],

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

E\mathbb{E}

Symbol E

The expected value operator: the probability-weighted average of the quantity inside its brackets.

Understand this part →

rθr_\theta

Symbol r_θ

r_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

xx

Symbol x

x is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

yy

Symbol y

y is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

β\beta

Symbol β

β is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

DKLD_{\mathrm{KL}}

Symbol D_KL

DKD_KL is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

πRL\pi^{\mathrm{RL}}

Symbol pi^RL

piRi^RL is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

πSFT\pi^{\mathrm{SFT}}

Symbol pi^SFT

piSi^SFT is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

γ\gamma

Symbol gamma

the pretraining mixture.

Understand this part →

addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subtraction

subtraction

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Policy optimisation. The fine-tuned model is then optimised, typically with PPO [ 6 ] , to maximise the reward model’s score, with a penalty on divergence from the starting policy. InstructGPT’s published objective adds a per-token KL penalty from the supervised model and, in the PPO-ptx variant, a term mixing in pretraining gradients: E[rθ(x,y)]−β E[DKL ⁣(πRL ∥ πSFT)]+γ E[log⁡πRL(x)]\mathbb{E}\big[r_\theta(x,y)\big] - \beta\,\mathbb{E}\big[D_{\mathrm{KL}}\!\left(\pi^{\mathrm{RL}} \,\|\, \pi^{\mathrm{SFT}}\right)\big] + \gamma\,\mathbb{E}\big[\log \pi^{\mathrm{RL}}(x)\big]. where β\beta sets the strength of the KL penalty and γ\gamma the pretraining mixture, with γ\gamma set to zero for the plain PPO models [ 4 ] . Direct preference optimisation later showed that the explicit reward model can be dispensed with entirely — the optimal policy under this objective has a closed form, so the same problem can be solved…
Read the full surrounding passage
Policy optimisation. The fine-tuned model is then optimised, typically with PPO [ 6 ] , to maximise the reward model’s score, with a penalty on divergence from the starting policy. InstructGPT’s published objective adds a per-token KL penalty from the supervised model and, in the PPO-ptx variant, a term mixing in pretraining gradients: E[rθ(x,y)]−β E[DKL ⁣(πRL ∥ πSFT)]+γ E[log⁡πRL(x)]\mathbb{E}\big[r_\theta(x,y)\big] - \beta\,\mathbb{E}\big[D_{\mathrm{KL}}\!\left(\pi^{\mathrm{RL}} \,\|\, \pi^{\mathrm{SFT}}\right)\big] + \gamma\,\mathbb{E}\big[\log \pi^{\mathrm{RL}}(x)\big]. where β\beta sets the strength of the KL penalty and γ\gamma the pretraining mixture, with γ\gamma set to zero for the plain PPO models [ 4 ] . Direct preference optimisation later showed that the explicit reward model can be dispensed with entirely — the optimal policy under this objective has a closed form, so the same problem can be solved with a classification loss on the preference pairs [ 15 ] . That is an important simplification, and it changes nothing about the present argument. DPO removes the reward network; it does not remove the reward. The objective is still rater approval, and the KL term is still there, folded into the loss.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness

See this formula across 1 published context →

Browse the mathematical compendium →