← All parts of this equation

Equation 6 · Part 11 · Claude, From First Principles: Training, Constitutional Methods, and What Actually Shapes a Response

=

J(θ)=Ex∼D, y∼πθ(⋅∣x)[r(x,y)]−βDKL(πθ(⋅∣x) ∥ πref(⋅∣x))J(\theta) = \mathbb{E}_{x \sim \mathcal{D},\ y \sim \pi_\theta(\cdot \mid x)}\left[r(x,y)\right] - \beta D_{\mathrm{KL}}\left(\pi_\theta(\cdot \mid x) \,\|\, \pi_{\mathrm{ref}}(\cdot \mid x)\right)
=

What this part means

The expressions on both sides represent the same quantity under the stated assumptions.

Its job in the formula

The equals sign connects the complete expression on the left with the complete expression on the right. Both sides must have compatible units.

The passage around this formula

The optimisation itself has a specific and revealing shape. Rather than maximising the fitted reward outright, the standard formulation constrains the updated policy to stay close to a reference policy, ordinarily the model just after supervised fine-tuning, one step removed from the raw pretrained prior: J(θ)=Ex∼D, y∼πθ(⋅∣x)[r(x,y)]−βDKL(πθ(⋅∣x) ∥ πref(⋅∣x))J(\theta) = \mathbb{E}_{x \sim \mathcal{D},\ y \sim \pi_\theta(\cdot \mid x)}\left[r(x,y)\right] - \beta D_{\mathrm{KL}}\left(\pi_\theta(\cdot \mid x) \,\|\, \pi_{\mathrm{ref}}(\cdot \mid x)\right). The reward term r(x,y) pulls the policy toward whatever the fitted preference model scores highly; the Kullback–Leibler penalty, weighted by β\beta , pulls it back toward the reference distribution. That second term is not a minor regulariser. It is the load-bearing assumption behind the entire method: remove it, and a policy optimised hard enough against an imperfect…

Read this part in the article →

Learn the underlying idea

An equals sign says that the expression on its left and the expression on its right have the same value under the stated definitions and assumptions.

Open the illustrated equality: what the equals sign claims guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.