← All parts of this equation

Equation 6 · Part 12 · Claude, From First Principles: Training, Constitutional Methods, and What Actually Shapes a Response

multiplication

J(θ)=Ex∼D, y∼πθ(⋅∣x)[r(x,y)]−βDKL(πθ(⋅∣x) ∥ πref(⋅∣x))J(\theta) = \mathbb{E}_{x \sim \mathcal{D},\ y \sim \pi_\theta(\cdot \mid x)}\left[r(x,y)\right] - \beta D_{\mathrm{KL}}\left(\pi_\theta(\cdot \mid x) \,\|\, \pi_{\mathrm{ref}}(\cdot \mid x)\right)
multiplication

What this part means

Multiply the quantities on either side.

Its job in the formula

Multiply the quantities on either side.

The passage around this formula

The optimisation itself has a specific and revealing shape. Rather than maximising the fitted reward outright, the standard formulation constrains the updated policy to stay close to a reference policy, ordinarily the model just after supervised fine-tuning, one step removed from the raw pretrained prior: J(θ)=Ex∼D, y∼πθ(⋅∣x)[r(x,y)]−βDKL(πθ(⋅∣x) ∥ πref(⋅∣x))J(\theta) = \mathbb{E}_{x \sim \mathcal{D},\ y \sim \pi_\theta(\cdot \mid x)}\left[r(x,y)\right] - \beta D_{\mathrm{KL}}\left(\pi_\theta(\cdot \mid x) \,\|\, \pi_{\mathrm{ref}}(\cdot \mid x)\right). The reward term r(x,y) pulls the policy toward whatever the fitted preference model scores highly; the Kullback–Leibler penalty, weighted by β\beta , pulls it back toward the reference distribution. That second term is not a minor regulariser. It is the load-bearing assumption behind the entire method: remove it, and a policy optimised hard enough against an imperfect…

Read this part in the article →

Learn the underlying idea

Multiplication scales one quantity by another. A dot, a cross, or adjacent symbols can indicate a product.

Open the illustrated multiplication: combining factors guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.