← All parts of this equation

Equation 11 · Part 4 · How Constitutional AI Actually Constrains a Model's Behavior

Symbol p

LPMAI(θ)=− E(x, yA, yB, p) ∼ DAI[p log⁡σ(rθ(x,yA)−rθ(x,yB))+(1−p) log⁡σ(rθ(x,yB)−rθ(x,yA))],\mathcal{L}^{AI}_{PM}(\theta) = -\,\mathbb{E}_{(x,\,y_A,\,y_B,\,p)\,\sim\, D_{AI}}\Big[p\,\log \sigma\big(r_\theta(x,y_A)-r_\theta(x,y_B)\big) + (1-p)\,\log \sigma\big(r_\theta(x,y_B)-r_\theta(x,y_A)\big)\Big],
pp

What this part means

the feedback model outputs a probability.

Its job in the formula

p appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Where the article explains it

For the soft-labeled AI comparisons specifically, where the feedback model outputs a probability p of preferring response yAy_A over yBy_B rather than a hard choice, the corresponding cross-entropy term is LPMAI(θ)=− E(x, yA, yB, p) ∼ DAI[p log⁡σ(rθ(x,yA)−rθ(x,yB))+(1−p) log⁡σ(rθ(x,yB)−rθ(x,yA))]\mathcal{L}^{AI}_{PM}(\theta) = -\,\mathbb{E}_{(x,\,y_A,\,y_B,\,p)\,\sim\, D_{AI}}\Big[p\,\log \sigma\big(r_\theta(x,y_A)-r_\theta(x,y_B)\big) + (1-p)\,\log \sigma\big(r_\theta(x,y_B)-r_\theta(x,y_A)\big)\Big].

The passage around this formula

…for one half of one dataset, not the loss, not the optimizer, not the use of a KL penalty against the supervised policy. For the soft-labeled AI comparisons specifically, where the feedback model outputs a probability p of preferring response yAy_A over yBy_B rather than a hard choice, the corresponding cross-entropy term is LPMAI(θ)=− E(x, yA, yB, p) ∼ DAI[p log⁡σ(rθ(x,yA)−rθ(x,yB))+(1−p) log⁡σ(rθ(x,yB)−rθ(x,yA))]\mathcal{L}^{AI}_{PM}(\theta) = -\,\mathbb{E}_{(x,\,y_A,\,y_B,\,p)\,\sim\, D_{AI}}\Big[p\,\log \sigma\big(r_\theta(x,y_A)-r_\theta(x,y_B)\big) + (1-p)\,\log \sigma\big(r_\theta(x,y_B)-r_\theta(x,y_A)\big)\Big]. which exposes the other genuine mechanical difference: AI-generated comparisons can carry a continuous confidence, where a human click is binary. Everything after the…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.