← All parts of this equation

Equation 2 · Part 10 · The Main Technical Approaches to AI Alignment, Compared

subscript

max⁡θ  Ex, y∼πθ[rϕ(x,y)]  −  β DKL(πθ ∥ πref).\max_\theta \; \mathbb{E}_{x,\, y\sim\pi_\theta}\big[r_\phi(x,y)\big] \;-\; \beta \, D_{\mathrm{KL}}\big(\pi_\theta \,\Vert\, \pi_{\mathrm{ref}}\big).
subscript

What this part means

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Its job in the formula

A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.

The passage around this formula

then optimizes the policy against that fitted reward under a penalty that keeps it near its starting point, max⁡θ  Ex, y∼πθ[rϕ(x,y)]  −  β DKL(πθ ∥ πref)\max_\theta \; \mathbb{E}_{x,\, y\sim\pi_\theta}\big[r_\phi(x,y)\big] \;-\; \beta \, D_{\mathrm{KL}}\big(\pi_\theta \,\Vert\, \pi_{\mathrm{ref}}\big). Read the two equations together and the oversight limit is visible in the structure itself, not only in later critiques of it. The reward model rϕr_\phi is trained on whatever a human rater could see and judge in the time given; nothing forces it to track truth or usefulness beyond what raters could distinguish. The KL term in the second equation is a standing admission of that limit — it constrains how far the policy may travel from the region where rϕr_\phi was actually calibrated, because past that region the fitted proxy is not trusted to mean anything. Casper and…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.