Equation 3 · The Main Technical Approaches to AI Alignment, Compared
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol r_phi
hi is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Read the two equations together and the oversight limit is visible in the structure itself, not only in later critiques of it. The reward model is trained on whatever a human rater could see and judge in the time given; nothing forces it to track truth or usefulness beyond what raters could distinguish. The KL term in the second equation is a standing admission of that limit — it constrains how far the policy may travel from the region where was actually calibrated, because past that region the fitted proxy is not trusted to mean anything. Casper and colleagues’ survey states the underlying identification directly: “reward models are trained to reflect human approval instead of…
Read the full surrounding passage
Read the two equations together and the oversight limit is visible in the structure itself, not only in later critiques of it. The reward model is trained on whatever a human rater could see and judge in the time given; nothing forces it to track truth or usefulness beyond what raters could distinguish. The KL term in the second equation is a standing admission of that limit — it constrains how far the policy may travel from the region where was actually calibrated, because past that region the fitted proxy is not trusted to mean anything. Casper and colleagues’ survey states the underlying identification directly: “reward models are trained to reflect human approval instead of human benefit,” and separately, that “a single reward function cannot represent a diverse society of humans,” concluding that condensing feedback from many people into one scalar is “a fundamentally misspecified problem” rather than a bug fixable within the RLHF framework [ 5 ] . Their paper is explicit that this is a distinct claim from RLHF’s more tractable engineering problems: a limitation counts as fundamental, in their terms, only when overcoming it “would require a method that is no longer a form of RLHF” [ 5 ] . RLHF’s scalable-oversight ceiling, in other words, is set by rater competence and by what a single scalar can represent — which is precisely the ceiling the other five techniques each try to raise or route around in a different way.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to The Main Technical Approaches to AI Alignment, Compared