← Back to article

Equation 11 · How Constitutional AI Actually Constrains a Model's Behavior

What does this equation mean?

LPMAI(θ)=− E(x, yA, yB, p) ∼ DAI[p log⁡σ(rθ(x,yA)−rθ(x,yB))+(1−p) log⁡σ(rθ(x,yB)−rθ(x,yA))],\mathcal{L}^{AI}_{PM}(\theta) = -\,\mathbb{E}_{(x,\,y_A,\,y_B,\,p)\,\sim\, D_{AI}}\Big[p\,\log \sigma\big(r_\theta(x,y_A)-r_\theta(x,y_B)\big) + (1-p)\,\log \sigma\big(r_\theta(x,y_B)-r_\theta(x,y_A)\big)\Big],

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

LPMAI\mathcal{L}^{AI}_{PM}

Symbol L^AI_PM

LAL^AIPI_PM is computed from the expected values combined on the right.

Understand this part →

θ\theta

Symbol θ

θ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

E(x, yA, yB, p) ∼ DAI\mathbb{E}_{(x,\,y_A,\,y_B,\,p)\,\sim\, D_{AI}}

Symbol E_(x,y_A,y_B,p)sim D_AI

E_(x,yAy_A,yBy_B,p)sim DAD_AI appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

pp

Symbol p

the feedback model outputs a probability.

Understand this part →

σ\sigma

Symbol σ

σ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

rθr_\theta

Symbol r_θ

r_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

xx

Symbol x

x appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

yAy_A

Symbol y_A

yAy_A appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

yBy_B

Symbol y_B

yBy_B appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Nothing in that loss distinguishes where a pair came from; a human-labeled winner and an AI-labeled winner are interchangeable once written down as (x, ywy_w, yly_l) . That is the precise, narrow sense in which Constitutional AI “differs mechanically from plain RLHF”: it changes the labeling function for one half of one dataset, not the loss, not the optimizer, not the use of a KL penalty against the supervised policy. For the soft-labeled AI comparisons specifically, where the feedback model outputs a probability p of preferring response yAy_A over yBy_B rather than a hard choice, the corresponding cross-entropy term is LPMAI(θ)=− E(x, yA, yB, p) ∼ DAI[p log⁡σ(rθ(x,yA)−rθ(x,yB))+(1−p) log⁡σ(rθ(x,yB)−rθ(x,yA))]\mathcal{L}^{AI}_{PM}(\theta) = -\,\mathbb{E}_{(x,\,y_A,\,y_B,\,p)\,\sim\, D_{AI}}\Big[p\,\log \sigma\big(r_\theta(x,y_A)-r_\theta(x,y_B)\big) + (1-p)\,\log \sigma\big(r_\theta(x,y_B)-r_\theta(x,y_A)\big)\Big]. which exposes the other genuine mechanical difference:…
Read the full surrounding passage
Nothing in that loss distinguishes where a pair came from; a human-labeled winner and an AI-labeled winner are interchangeable once written down as (x, ywy_w, yly_l) . That is the precise, narrow sense in which Constitutional AI “differs mechanically from plain RLHF”: it changes the labeling function for one half of one dataset, not the loss, not the optimizer, not the use of a KL penalty against the supervised policy. For the soft-labeled AI comparisons specifically, where the feedback model outputs a probability p of preferring response yAy_A over yBy_B rather than a hard choice, the corresponding cross-entropy term is LPMAI(θ)=− E(x, yA, yB, p) ∼ DAI[p log⁡σ(rθ(x,yA)−rθ(x,yB))+(1−p) log⁡σ(rθ(x,yB)−rθ(x,yA))]\mathcal{L}^{AI}_{PM}(\theta) = -\,\mathbb{E}_{(x,\,y_A,\,y_B,\,p)\,\sim\, D_{AI}}\Big[p\,\log \sigma\big(r_\theta(x,y_A)-r_\theta(x,y_B)\big) + (1-p)\,\log \sigma\big(r_\theta(x,y_B)-r_\theta(x,y_A)\big)\Big]. which exposes the other genuine mechanical difference: AI-generated comparisons can carry a continuous confidence, where a human click is binary. Everything after the preference model, the PPO update against it with a KL penalty toward the SL-CAI policy, is the RLHF recipe unmodified.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to How Constitutional AI Actually Constrains a Model's Behavior

See this formula across 1 published context →

Browse the mathematical compendium →