← All parts of this equation

Equation 6 · Part 13 · Claude, From First Principles: Training, Constitutional Methods, and What Actually Shapes a Response

subtraction

J(θ)=Ex∼D, y∼πθ(⋅∣x)[r(x,y)]−βDKL(πθ(⋅∣x) ∥ πref(⋅∣x))J(\theta) = \mathbb{E}_{x \sim \mathcal{D},\ y \sim \pi_\theta(\cdot \mid x)}\left[r(x,y)\right] - \beta D_{\mathrm{KL}}\left(\pi_\theta(\cdot \mid x) \,\|\, \pi_{\mathrm{ref}}(\cdot \mid x)\right)
subtraction

What this part means

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Its job in the formula

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

The passage around this formula

The optimisation itself has a specific and revealing shape. Rather than maximising the fitted reward outright, the standard formulation constrains the updated policy to stay close to a reference policy, ordinarily the model just after supervised fine-tuning, one step removed from the raw pretrained prior: J(θ)=Ex∼D, y∼πθ(⋅∣x)[r(x,y)]−βDKL(πθ(⋅∣x) ∥ πref(⋅∣x))J(\theta) = \mathbb{E}_{x \sim \mathcal{D},\ y \sim \pi_\theta(\cdot \mid x)}\left[r(x,y)\right] - \beta D_{\mathrm{KL}}\left(\pi_\theta(\cdot \mid x) \,\|\, \pi_{\mathrm{ref}}(\cdot \mid x)\right). The reward term r(x,y) pulls the policy toward whatever the fitted preference model scores highly; the Kullback–Leibler penalty, weighted by β\beta , pulls it back toward the reference distribution. That second term is not a minor regulariser. It is the load-bearing assumption behind the entire method: remove it, and a policy optimised hard enough against an imperfect…

Read this part in the article →

Learn the underlying idea

Addition combines quantities; subtraction measures the signed difference between them. Parentheses show what is combined before the rest of the expression is evaluated.

Open the illustrated addition and subtraction in an equation guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.