Published equation contexts
Why this formula appears here
The optimisation itself has a specific and revealing shape. Rather than maximising the fitted reward outright, the standard formulation constrains the updated policy to stay close to a reference policy, ordinarily the model just after supervised fine-tuning, one step removed from the raw pretrained prior: . The reward term r(x,y) pulls the policy toward whatever the fitted preference model scores highly; the Kullback–Leibler penalty, weighted by , pulls it back toward the reference distribution. That second term is not a minor regulariser. It is the load-bearing assumption behind the entire method: remove it, and a policy optimised hard enough against an imperfect…
Read the representative guide
Symbol θ
θ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Read this term in its guide →Symbol E_x sim D, y sim pi_θ( × mid x)
sim D, y sim pi_θ( × mid x) appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →Symbol r
r appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →Symbol x
x appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →Symbol y
y appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →Symbol D_KL
L appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →Symbol pi_θ
pi_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →Symbol pi_ref
pef appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
Research cited beside this formula
Published contexts (1)
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 6 · Foundation Models
Claude, From First Principles: Training, Constitutional Methods, and What Actually Shapes a Response
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
The optimisation itself has a specific and revealing shape. Rather than maximising the fitted reward outright, the standard formulation constrains the updated policy to stay close to a reference policy, ordinarily the model just after supervised fine-tuning, one step removed from the raw pretrained prior: . The reward term r(x,y) pulls the policy toward whatever the fitted preference model scores highly; the Kullback–Leibler penalty, weighted by , pulls it back toward the reference distribution. That second term is not a minor regulariser. It is the load-bearing assumption behind the entire method: remove it, and a policy optimised hard enough against an imperfect…