Equation 5 · Part 6 · The Main Technical Approaches to AI Alignment, Compared
Symbol σ
What this part means
σ is one of the signed contributions combined to compute the quantity on the left.
Its job in the formula
σ is one of the signed contributions combined to compute the quantity on the left.
Full expression→Symbol σ→Article meaning
The passage around this formula
Constitutional AI keeps RLHF’s reward-model-and-KL-penalty backbone intact and changes where the comparison labels come from. Bai and colleagues describe a two-phase method: a supervised phase in which the model critiques and revises its own responses against a written set of principles, and a reinforcement phase in which a model, rather than a human, judges which of two candidate responses better satisfies those principles — producing an AI-generated preference dataset that trains the reward model [ 6 ] . Formally, this changes only the source of the comparison label. The Bradley–Terry equation above is unchanged in form; what changes is that the probability being fitted is now [displayed…
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the surrounding passage
- [6] Constitutional AI: Harmlessness from AI Feedback ↗
- [7] RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback ↗
These citations provide research context; check each source for the exact claim it supports.