← Back to article

Equation 4 · The Main Technical Approaches to AI Alignment, Compared

What does this equation mean?

rϕr_\phi

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

rϕr_\phi

Symbol r_phi

rpr_phi is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Read the two equations together and the oversight limit is visible in the structure itself, not only in later critiques of it. The reward model rϕr_\phi is trained on whatever a human rater could see and judge in the time given; nothing forces it to track truth or usefulness beyond what raters could distinguish. The KL term in the second equation is a standing admission of that limit — it constrains how far the policy may travel from the region where rϕr_\phi was actually calibrated, because past that region the fitted proxy is not trusted to mean anything. Casper and colleagues’ survey states the underlying identification directly: “reward models are trained to reflect human approval instead of…
Read the full surrounding passage
Read the two equations together and the oversight limit is visible in the structure itself, not only in later critiques of it. The reward model rϕr_\phi is trained on whatever a human rater could see and judge in the time given; nothing forces it to track truth or usefulness beyond what raters could distinguish. The KL term in the second equation is a standing admission of that limit — it constrains how far the policy may travel from the region where rϕr_\phi was actually calibrated, because past that region the fitted proxy is not trusted to mean anything. Casper and colleagues’ survey states the underlying identification directly: “reward models are trained to reflect human approval instead of human benefit,” and separately, that “a single reward function cannot represent a diverse society of humans,” concluding that condensing feedback from many people into one scalar is “a fundamentally misspecified problem” rather than a bug fixable within the RLHF framework [ 5 ] . Their paper is explicit that this is a distinct claim from RLHF’s more tractable engineering problems: a limitation counts as fundamental, in their terms, only when overcoming it “would require a method that is no longer a form of RLHF” [ 5 ] . RLHF’s scalable-oversight ceiling, in other words, is set by rater competence and by what a single scalar can represent — which is precisely the ceiling the other five techniques each try to raise or route around in a different way.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to The Main Technical Approaches to AI Alignment, Compared

Browse the mathematical compendium →