← All parts of this equation

Equation 2 · Part 8 · The Main Technical Approaches to AI Alignment, Compared

Symbol pi_θ

max⁡θ  Ex, y∼πθ[rϕ(x,y)]  −  β DKL(πθ ∥ πref).\max_\theta \; \mathbb{E}_{x,\, y\sim\pi_\theta}\big[r_\phi(x,y)\big] \;-\; \beta \, D_{\mathrm{KL}}\big(\pi_\theta \,\Vert\, \pi_{\mathrm{ref}}\big).
πθ\pi_\theta

What this part means

pi_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Its job in the formula

pi_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

The passage around this formula

then optimizes the policy against that fitted reward under a penalty that keeps it near its starting point, max⁡θ  Ex, y∼πθ[rϕ(x,y)]  −  β DKL(πθ ∥ πref)\max_\theta \; \mathbb{E}_{x,\, y\sim\pi_\theta}\big[r_\phi(x,y)\big] \;-\; \beta \, D_{\mathrm{KL}}\big(\pi_\theta \,\Vert\, \pi_{\mathrm{ref}}\big). Read the two equations together and the oversight limit is visible in the structure itself, not only in later critiques of it. The reward model rϕr_\phi is trained on whatever a human rater could see and judge in the time given; nothing forces it to track truth or usefulness beyond what raters could distinguish. The KL term in the second equation is a standing admission of that limit — it constrains how far the policy may travel from the region where rϕr_\phi was actually calibrated, because past that region the fitted proxy is not trusted to mean anything. Casper and…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.