Equation 6 · How Constitutional AI Actually Constrains a Model's Behavior
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol θ
θ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Symbol E_(x,y_w,y_l)sim D_H cup D_AI
E_(x,,)sim cup I appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol σ
σ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol r_θ
r_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol x
x appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol y_w
appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol y_l
appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Crucially, this AI-generated harmlessness comparison data did not replace human comparison data outright; it was combined with it. The paper reports 135,296 human helpfulness comparisons and 182,831 constitutionally generated harmlessness comparisons feeding one preference model, trained on the union of both, in the authors’ words: “we use human labels for helpfulness, but only AI labels for harmlessness.” A single Bradley-Terry-style preference loss is fit across both sources at once. Writing for the human comparison set, for the AI-generated set, for the reward model, x for a prompt, and , for the winning and losing response in a pair, the objective is [displayed…
Read the full surrounding passage
Crucially, this AI-generated harmlessness comparison data did not replace human comparison data outright; it was combined with it. The paper reports 135,296 human helpfulness comparisons and 182,831 constitutionally generated harmlessness comparisons feeding one preference model, trained on the union of both, in the authors’ words: “we use human labels for helpfulness, but only AI labels for harmlessness.” A single Bradley-Terry-style preference loss is fit across both sources at once. Writing for the human comparison set, for the AI-generated set, for the reward model, x for a prompt, and , for the winning and losing response in a pair, the objective is . Nothing in that loss distinguishes where a pair came from; a human-labeled winner and an AI-labeled winner are interchangeable once written down as (x, , ) . That is the precise, narrow sense in which Constitutional AI “differs mechanically from plain RLHF”: it changes the labeling function for one half of one dataset, not the loss, not the optimizer, not the use of a KL penalty against the supervised policy. For the soft-labeled AI comparisons specifically, where the feedback model outputs a probability p of preferring response over rather than a hard choice, the corresponding cross-entropy term is
Sources cited in the article section
These citations give research context. Read each source to check which claims it supports.
Return to How Constitutional AI Actually Constrains a Model's Behavior