Equation 5 · Part 8 · The Main Technical Approaches to AI Alignment, Compared
=
=
What this part means
The expressions on both sides represent the same quantity under the stated assumptions.
Its job in the formula
The equals sign connects the complete expression on the left with the complete expression on the right. Both sides must have compatible units.
Full expression→=→Article meaning
The passage around this formula
Constitutional AI keeps RLHF’s reward-model-and-KL-penalty backbone intact and changes where the comparison labels come from. Bai and colleagues describe a two-phase method: a supervised phase in which the model critiques and revises its own responses against a written set of principles, and a reinforcement phase in which a model, rather than a human, judges which of two candidate responses better satisfies those principles — producing an AI-generated preference dataset that trains the reward model [ 6 ] . Formally, this changes only the source of the comparison label. The Bradley–Terry equation above is unchanged in form; what changes is that the probability being fitted is now [displayed…
Learn the underlying idea
An equals sign says that the expression on its left and the expression on its right have the same value under the stated definitions and assumptions.
Open the illustrated equality: what the equals sign claims guide →
Sources cited in the surrounding passage
- [6] Constitutional AI: Harmlessness from AI Feedback ↗
- [7] RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback ↗
These citations provide research context; check each source for the exact claim it supports.