← Mathematical compendium

Published equation contexts

J(θ)=Ex∼D, y∼πθ(⋅∣x)[r(x,y)]−βDKL(πθ(⋅∣x) ∥ πref(⋅∣x))J(\theta) = \mathbb{E}_{x \sim \mathcal{D},\ y \sim \pi_\theta(\cdot \mid x)}\left[r(x,y)\right] - \beta D_{\mathrm{KL}}\left(\pi_\theta(\cdot \mid x) \,\|\, \pi_{\mathrm{ref}}(\cdot \mid x)\right)

Why this formula appears here

The optimisation itself has a specific and revealing shape. Rather than maximising the fitted reward outright, the standard formulation constrains the updated policy to stay close to a reference policy, ordinarily the model just after supervised fine-tuning, one step removed from the raw pretrained prior: J(θ)=Ex∼D, y∼πθ(⋅∣x)[r(x,y)]−βDKL(πθ(⋅∣x) ∥ πref(⋅∣x))J(\theta) = \mathbb{E}_{x \sim \mathcal{D},\ y \sim \pi_\theta(\cdot \mid x)}\left[r(x,y)\right] - \beta D_{\mathrm{KL}}\left(\pi_\theta(\cdot \mid x) \,\|\, \pi_{\mathrm{ref}}(\cdot \mid x)\right). The reward term r(x,y) pulls the policy toward whatever the fitted preference model scores highly; the Kullback–Leibler penalty, weighted by β\beta , pulls it back toward the reference distribution. That second term is not a minor regulariser. It is the load-bearing assumption behind the entire method: remove it, and a policy optimised hard enough against an imperfect…

Read the full article-specific guide →

Read the representative guide

θ\theta

Symbol θ

θ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Read this term in its guide →
Ex∼D, y∼πθ(⋅∣x)\mathbb{E}_{x \sim \mathcal{D},\ y \sim \pi_\theta(\cdot \mid x)}

Symbol E_x sim D, y sim pi_θ( × mid x)

ExE_x sim D, y sim pi_θ( × mid x) appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
rr

Symbol r

r appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
xx

Symbol x

x appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
yy

Symbol y

y appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
DKLD_{\mathrm{KL}}

Symbol D_KL

DKD_KL appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
πθ\pi_\theta

Symbol pi_θ

pi_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
πref\pi_{\mathrm{ref}}

Symbol pi_ref

piri_ref appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

J(θ)=Ex∼D, y∼πθ(⋅∣x)[r(x,y)]−βDKL(πθ(⋅∣x) ∥ πref(⋅∣x))J(\theta) = \mathbb{E}_{x \sim \mathcal{D},\ y \sim \pi_\theta(\cdot \mid x)}\left[r(x,y)\right] - \beta D_{\mathrm{KL}}\left(\pi_\theta(\cdot \mid x) \,\|\, \pi_{\mathrm{ref}}(\cdot \mid x)\right)

Equation 6 · Foundation Models

Claude, From First Principles: Training, Constitutional Methods, and What Actually Shapes a Response

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

The optimisation itself has a specific and revealing shape. Rather than maximising the fitted reward outright, the standard formulation constrains the updated policy to stay close to a reference policy, ordinarily the model just after supervised fine-tuning, one step removed from the raw pretrained prior: J(θ)=Ex∼D, y∼πθ(⋅∣x)[r(x,y)]−βDKL(πθ(⋅∣x) ∥ πref(⋅∣x))J(\theta) = \mathbb{E}_{x \sim \mathcal{D},\ y \sim \pi_\theta(\cdot \mid x)}\left[r(x,y)\right] - \beta D_{\mathrm{KL}}\left(\pi_\theta(\cdot \mid x) \,\|\, \pi_{\mathrm{ref}}(\cdot \mid x)\right). The reward term r(x,y) pulls the policy toward whatever the fitted preference model scores highly; the Kullback–Leibler penalty, weighted by β\beta , pulls it back toward the reference distribution. That second term is not a minor regulariser. It is the load-bearing assumption behind the entire method: remove it, and a policy optimised hard enough against an imperfect…

Meanings in this article

Equation guide → · Article →