← Back to article

Equation 2 · The Main Technical Approaches to AI Alignment, Compared

What does this equation mean?

max⁡θ  Ex, y∼πθ[rϕ(x,y)]  −  β DKL(πθ ∥ πref).\max_\theta \; \mathbb{E}_{x,\, y\sim\pi_\theta}\big[r_\phi(x,y)\big] \;-\; \beta \, D_{\mathrm{KL}}\big(\pi_\theta \,\Vert\, \pi_{\mathrm{ref}}\big).

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

θ\theta

Symbol θ

θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

Ex, y∼πθ\mathbb{E}_{x,\, y\sim\pi_\theta}

Symbol E_x, ysimpi_θ

ExE_x, ysimpi_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

rϕr_\phi

Symbol r_phi

rpr_phi is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

xx

Symbol x

x is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

yy

Symbol y

y is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

β\beta

Symbol β

β is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

DKLD_{\mathrm{KL}}

Symbol D_KL

DKD_KL is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

πθ\pi_\theta

Symbol pi_θ

pi_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

πref\pi_{\mathrm{ref}}

Symbol pi_ref

piri_ref is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

then optimizes the policy against that fitted reward under a penalty that keeps it near its starting point, max⁡θ  Ex, y∼πθ[rϕ(x,y)]  −  β DKL(πθ ∥ πref)\max_\theta \; \mathbb{E}_{x,\, y\sim\pi_\theta}\big[r_\phi(x,y)\big] \;-\; \beta \, D_{\mathrm{KL}}\big(\pi_\theta \,\Vert\, \pi_{\mathrm{ref}}\big). Read the two equations together and the oversight limit is visible in the structure itself, not only in later critiques of it. The reward model rϕr_\phi is trained on whatever a human rater could see and judge in the time given; nothing forces it to track truth or usefulness beyond what raters could distinguish. The KL term in the second equation is a standing admission of that limit — it constrains how far the policy may travel from the region where rϕr_\phi was actually calibrated, because past that region the fitted proxy is not trusted to mean anything. Casper and…
Read the full surrounding passage
then optimizes the policy against that fitted reward under a penalty that keeps it near its starting point, max⁡θ  Ex, y∼πθ[rϕ(x,y)]  −  β DKL(πθ ∥ πref)\max_\theta \; \mathbb{E}_{x,\, y\sim\pi_\theta}\big[r_\phi(x,y)\big] \;-\; \beta \, D_{\mathrm{KL}}\big(\pi_\theta \,\Vert\, \pi_{\mathrm{ref}}\big). Read the two equations together and the oversight limit is visible in the structure itself, not only in later critiques of it. The reward model rϕr_\phi is trained on whatever a human rater could see and judge in the time given; nothing forces it to track truth or usefulness beyond what raters could distinguish. The KL term in the second equation is a standing admission of that limit — it constrains how far the policy may travel from the region where rϕr_\phi was actually calibrated, because past that region the fitted proxy is not trusted to mean anything. Casper and colleagues’ survey states the underlying identification directly: “reward models are trained to reflect human approval instead of human benefit,” and separately, that “a single reward function cannot represent a diverse society of humans,” concluding that condensing feedback from many people into one scalar is “a fundamentally misspecified problem” rather than a bug fixable within the RLHF framework [ 5 ] . Their paper is explicit that this is a distinct claim from RLHF’s more tractable engineering problems: a limitation counts as fundamental, in their terms, only when overcoming it “would require a method that is no longer a form of RLHF” [ 5 ] . RLHF’s scalable-oversight ceiling, in other words, is set by rater competence and by what a single scalar can represent — which is precisely the ceiling the other five techniques each try to raise or route around in a different way.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to The Main Technical Approaches to AI Alignment, Compared

See this formula across 1 published context →

Browse the mathematical compendium →