← All parts of this equation

Equation 3 · Part 12 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness

subscript

E[rθ(x,y)]−β E[DKL ⁣(πRL ∥ πSFT)]+γ E[log⁡πRL(x)],\mathbb{E}\big[r_\theta(x,y)\big] - \beta\,\mathbb{E}\big[D_{\mathrm{KL}}\!\left(\pi^{\mathrm{RL}} \,\|\, \pi^{\mathrm{SFT}}\right)\big] + \gamma\,\mathbb{E}\big[\log \pi^{\mathrm{RL}}(x)\big],
subscript

What this part means

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Its job in the formula

A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.

The passage around this formula

Policy optimisation. The fine-tuned model is then optimised, typically with PPO [ 6 ] , to maximise the reward model’s score, with a penalty on divergence from the starting policy. InstructGPT’s published objective adds a per-token KL penalty from the supervised model and, in the PPO-ptx variant, a term mixing in pretraining gradients: E[rθ(x,y)]−β E[DKL ⁣(πRL ∥ πSFT)]+γ E[log⁡πRL(x)]\mathbb{E}\big[r_\theta(x,y)\big] - \beta\,\mathbb{E}\big[D_{\mathrm{KL}}\!\left(\pi^{\mathrm{RL}} \,\|\, \pi^{\mathrm{SFT}}\right)\big] + \gamma\,\mathbb{E}\big[\log \pi^{\mathrm{RL}}(x)\big]. where β\beta sets the strength of the KL penalty and γ\gamma the pretraining mixture, with γ\gamma set to zero for the plain PPO models [ 4 ] . Direct preference optimisation later showed that the explicit reward model can be dispensed with entirely — the optimal policy under this objective has a closed form, so the same problem can be solved…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.