← Mathematical compendium

Published equation contexts

LPMAI(θ)=− E(x, yA, yB, p) ∼ DAI[p log⁡σ(rθ(x,yA)−rθ(x,yB))+(1−p) log⁡σ(rθ(x,yB)−rθ(x,yA))]\mathcal{L}^{AI}_{PM}(\theta) = -\,\mathbb{E}_{(x,\,y_A,\,y_B,\,p)\,\sim\, D_{AI}}\Big[p\,\log \sigma\big(r_\theta(x,y_A)-r_\theta(x,y_B)\big) + (1-p)\,\log \sigma\big(r_\theta(x,y_B)-r_\theta(x,y_A)\big)\Big]

Why this formula appears here

Nothing in that loss distinguishes where a pair came from; a human-labeled winner and an AI-labeled winner are interchangeable once written down as (x, ywy_w, yly_l) . That is the precise, narrow sense in which Constitutional AI “differs mechanically from plain RLHF”: it changes the labeling function for one half of one dataset, not the loss, not the optimizer, not the use of a KL penalty against the supervised policy. For the soft-labeled AI comparisons specifically, where the feedback model outputs a probability p of preferring response yAy_A over yBy_B rather than a hard choice, the corresponding cross-entropy term is LPMAI(θ)=− E(x, yA, yB, p) ∼ DAI[p log⁡σ(rθ(x,yA)−rθ(x,yB))+(1−p) log⁡σ(rθ(x,yB)−rθ(x,yA))]\mathcal{L}^{AI}_{PM}(\theta) = -\,\mathbb{E}_{(x,\,y_A,\,y_B,\,p)\,\sim\, D_{AI}}\Big[p\,\log \sigma\big(r_\theta(x,y_A)-r_\theta(x,y_B)\big) + (1-p)\,\log \sigma\big(r_\theta(x,y_B)-r_\theta(x,y_A)\big)\Big]. which exposes the other genuine mechanical difference:…

Read the full article-specific guide →

Read the representative guide

LPMAI\mathcal{L}^{AI}_{PM}

Symbol L^AI_PM

LAL^AIPI_PM is computed from the expected values combined on the right.

Read this term in its guide →
θ\theta

Symbol θ

θ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Read this term in its guide →
E(x, yA, yB, p) ∼ DAI\mathbb{E}_{(x,\,y_A,\,y_B,\,p)\,\sim\, D_{AI}}

Symbol E_(x,y_A,y_B,p)sim D_AI

E_(x,yAy_A,yBy_B,p)sim DAD_AI appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
σ\sigma

Symbol σ

σ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
rθr_\theta

Symbol r_θ

r_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
xx

Symbol x

x appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
yAy_A

Symbol y_A

yAy_A appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
yBy_B

Symbol y_B

yBy_B appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

LPMAI(θ)=− E(x, yA, yB, p) ∼ DAI[p log⁡σ(rθ(x,yA)−rθ(x,yB))+(1−p) log⁡σ(rθ(x,yB)−rθ(x,yA))],\mathcal{L}^{AI}_{PM}(\theta) = -\,\mathbb{E}_{(x,\,y_A,\,y_B,\,p)\,\sim\, D_{AI}}\Big[p\,\log \sigma\big(r_\theta(x,y_A)-r_\theta(x,y_B)\big) + (1-p)\,\log \sigma\big(r_\theta(x,y_B)-r_\theta(x,y_A)\big)\Big],

Equation 11 · Foundation Models

How Constitutional AI Actually Constrains a Model's Behavior

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Nothing in that loss distinguishes where a pair came from; a human-labeled winner and an AI-labeled winner are interchangeable once written down as (x, ywy_w, yly_l) . That is the precise, narrow sense in which Constitutional AI “differs mechanically from plain RLHF”: it changes the labeling function for one half of one dataset, not the loss, not the optimizer, not the use of a KL penalty against the supervised policy. For the soft-labeled AI comparisons specifically, where the feedback model outputs a probability p of preferring response yAy_A over yBy_B rather than a hard choice, the corresponding cross-entropy term is LPMAI(θ)=− E(x, yA, yB, p) ∼ DAI[p log⁡σ(rθ(x,yA)−rθ(x,yB))+(1−p) log⁡σ(rθ(x,yB)−rθ(x,yA))]\mathcal{L}^{AI}_{PM}(\theta) = -\,\mathbb{E}_{(x,\,y_A,\,y_B,\,p)\,\sim\, D_{AI}}\Big[p\,\log \sigma\big(r_\theta(x,y_A)-r_\theta(x,y_B)\big) + (1-p)\,\log \sigma\big(r_\theta(x,y_B)-r_\theta(x,y_A)\big)\Big]. which exposes the other genuine mechanical difference:…

Meanings in this article

  • pp: the feedback model outputs a probability.
Equation guide → · Article →