Equation 11 · How Constitutional AI Actually Constrains a Model's Behavior
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol θ
θ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Symbol E_(x,y_A,y_B,p)sim D_AI
E_(x,,,p)sim I appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol σ
σ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol r_θ
r_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol x
x appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol y_A
appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol y_B
appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Nothing in that loss distinguishes where a pair came from; a human-labeled winner and an AI-labeled winner are interchangeable once written down as (x, , ) . That is the precise, narrow sense in which Constitutional AI “differs mechanically from plain RLHF”: it changes the labeling function for one half of one dataset, not the loss, not the optimizer, not the use of a KL penalty against the supervised policy. For the soft-labeled AI comparisons specifically, where the feedback model outputs a probability p of preferring response over rather than a hard choice, the corresponding cross-entropy term is . which exposes the other genuine mechanical difference:…
Read the full surrounding passage
Nothing in that loss distinguishes where a pair came from; a human-labeled winner and an AI-labeled winner are interchangeable once written down as (x, , ) . That is the precise, narrow sense in which Constitutional AI “differs mechanically from plain RLHF”: it changes the labeling function for one half of one dataset, not the loss, not the optimizer, not the use of a KL penalty against the supervised policy. For the soft-labeled AI comparisons specifically, where the feedback model outputs a probability p of preferring response over rather than a hard choice, the corresponding cross-entropy term is . which exposes the other genuine mechanical difference: AI-generated comparisons can carry a continuous confidence, where a human click is binary. Everything after the preference model, the PPO update against it with a KL penalty toward the SL-CAI policy, is the RLHF recipe unmodified.
Sources cited in the article section
These citations give research context. Read each source to check which claims it supports.
Return to How Constitutional AI Actually Constrains a Model's Behavior