← Mathematical compendium

Published equation contexts

max⁡θ  Ex, y∼πθ[rϕ(x,y)]  −  β DKL(πθ ∥ πref)\max_\theta \; \mathbb{E}_{x,\, y\sim\pi_\theta}\big[r_\phi(x,y)\big] \;-\; \beta \, D_{\mathrm{KL}}\big(\pi_\theta \,\Vert\, \pi_{\mathrm{ref}}\big)

Why this formula appears here

then optimizes the policy against that fitted reward under a penalty that keeps it near its starting point, max⁡θ  Ex, y∼πθ[rϕ(x,y)]  −  β DKL(πθ ∥ πref)\max_\theta \; \mathbb{E}_{x,\, y\sim\pi_\theta}\big[r_\phi(x,y)\big] \;-\; \beta \, D_{\mathrm{KL}}\big(\pi_\theta \,\Vert\, \pi_{\mathrm{ref}}\big). Read the two equations together and the oversight limit is visible in the structure itself, not only in later critiques of it. The reward model rϕr_\phi is trained on whatever a human rater could see and judge in the time given; nothing forces it to track truth or usefulness beyond what raters could distinguish. The KL term in the second equation is a standing admission of that limit — it constrains how far the policy may travel from the region where rϕr_\phi was actually calibrated, because past that region the fitted proxy is not trusted to mean anything. Casper and…

Read the full article-specific guide →

Read the representative guide

θ\theta

Symbol θ

θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
Ex, y∼πθ\mathbb{E}_{x,\, y\sim\pi_\theta}

Symbol E_x, ysimpi_θ

ExE_x, ysimpi_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
rϕr_\phi

Symbol r_phi

rpr_phi is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
xx

Symbol x

x is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
yy

Symbol y

y is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
β\beta

Symbol β

β is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
DKLD_{\mathrm{KL}}

Symbol D_KL

DKD_KL is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
πθ\pi_\theta

Symbol pi_θ

pi_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
πref\pi_{\mathrm{ref}}

Symbol pi_ref

piri_ref is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

max⁡θ  Ex, y∼πθ[rϕ(x,y)]  −  β DKL(πθ ∥ πref).\max_\theta \; \mathbb{E}_{x,\, y\sim\pi_\theta}\big[r_\phi(x,y)\big] \;-\; \beta \, D_{\mathrm{KL}}\big(\pi_\theta \,\Vert\, \pi_{\mathrm{ref}}\big).

Equation 2 · AI Safety

The Main Technical Approaches to AI Alignment, Compared

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

then optimizes the policy against that fitted reward under a penalty that keeps it near its starting point, max⁡θ  Ex, y∼πθ[rϕ(x,y)]  −  β DKL(πθ ∥ πref)\max_\theta \; \mathbb{E}_{x,\, y\sim\pi_\theta}\big[r_\phi(x,y)\big] \;-\; \beta \, D_{\mathrm{KL}}\big(\pi_\theta \,\Vert\, \pi_{\mathrm{ref}}\big). Read the two equations together and the oversight limit is visible in the structure itself, not only in later critiques of it. The reward model rϕr_\phi is trained on whatever a human rater could see and judge in the time given; nothing forces it to track truth or usefulness beyond what raters could distinguish. The KL term in the second equation is a standing admission of that limit — it constrains how far the policy may travel from the region where rϕr_\phi was actually calibrated, because past that region the fitted proxy is not trusted to mean anything. Casper and…

Equation guide → · Article →