Symbol E
The expected value operator: the probability-weighted average of the quantity inside its brackets.
Read this term in its guide →Published equation contexts
Policy optimisation. The fine-tuned model is then optimised, typically with PPO [ 6 ] , to maximise the reward model’s score, with a penalty on divergence from the starting policy. InstructGPT’s published objective adds a per-token KL penalty from the supervised model and, in the PPO-ptx variant, a term mixing in pretraining gradients: . where sets the strength of the KL penalty and the pretraining mixture, with set to zero for the plain PPO models [ 4 ] . Direct preference optimisation later showed that the explicit reward model can be dispensed with entirely — the optimal policy under this objective has a closed form, so the same problem can be solved…
The expected value operator: the probability-weighted average of the quantity inside its brackets.
Read this term in its guide →r_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →x is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →y is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →β is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →L is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →pL is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →pFT is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Read this expression with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 3 · Alignment & Safety
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
Policy optimisation. The fine-tuned model is then optimised, typically with PPO [ 6 ] , to maximise the reward model’s score, with a penalty on divergence from the starting policy. InstructGPT’s published objective adds a per-token KL penalty from the supervised model and, in the PPO-ptx variant, a term mixing in pretraining gradients: . where sets the strength of the KL penalty and the pretraining mixture, with set to zero for the plain PPO models [ 4 ] . Direct preference optimisation later showed that the explicit reward model can be dispensed with entirely — the optimal policy under this objective has a closed form, so the same problem can be solved…