← Mathematical compendium

Published equation contexts

LPM(θ)=− E(x, yw, yl) ∼ DH∪DAI[log⁡σ(rθ(x,yw)−rθ(x,yl))]\mathcal{L}_{PM}(\theta) = -\,\mathbb{E}_{(x,\,y_w,\,y_l)\,\sim\, D_H \cup D_{AI}}\Big[\log \sigma\big(r_\theta(x,y_w) - r_\theta(x,y_l)\big)\Big]

Why this formula appears here

Crucially, this AI-generated harmlessness comparison data did not replace human comparison data outright; it was combined with it. The paper reports 135,296 human helpfulness comparisons and 182,831 constitutionally generated harmlessness comparisons feeding one preference model, trained on the union of both, in the authors’ words: “we use human labels for helpfulness, but only AI labels for harmlessness.” A single Bradley-Terry-style preference loss is fit across both sources at once. Writing DHD_H for the human comparison set, DAID_{AI} for the AI-generated set, rθr_\theta for the reward model, x for a prompt, and ywy_w, yly_l for the winning and losing response in a pair, the objective is [displayed…

Read the full article-specific guide →

Read the representative guide

θ\theta

Symbol θ

θ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Read this term in its guide →
E(x, yw, yl) ∼ DH∪DAI\mathbb{E}_{(x,\,y_w,\,y_l)\,\sim\, D_H \cup D_{AI}}

Symbol E_(x,y_w,y_l)sim D_H cup D_AI

E_(x,ywy_w,yly_l)sim DHD_H cup DAD_AI appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
σ\sigma

Symbol σ

σ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
rθr_\theta

Symbol r_θ

r_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
xx

Symbol x

x appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
ywy_w

Symbol y_w

ywy_w appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
yly_l

Symbol y_l

yly_l appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

LPM(θ)=− E(x, yw, yl) ∼ DH∪DAI[log⁡σ(rθ(x,yw)−rθ(x,yl))].\mathcal{L}_{PM}(\theta) = -\,\mathbb{E}_{(x,\,y_w,\,y_l)\,\sim\, D_H \cup D_{AI}}\Big[\log \sigma\big(r_\theta(x,y_w) - r_\theta(x,y_l)\big)\Big].

Equation 6 · Foundation Models

How Constitutional AI Actually Constrains a Model's Behavior

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Crucially, this AI-generated harmlessness comparison data did not replace human comparison data outright; it was combined with it. The paper reports 135,296 human helpfulness comparisons and 182,831 constitutionally generated harmlessness comparisons feeding one preference model, trained on the union of both, in the authors’ words: “we use human labels for helpfulness, but only AI labels for harmlessness.” A single Bradley-Terry-style preference loss is fit across both sources at once. Writing DHD_H for the human comparison set, DAID_{AI} for the AI-generated set, rθr_\theta for the reward model, x for a prompt, and ywy_w, yly_l for the winning and losing response in a pair, the objective is [displayed…

Equation guide → · Article →