← All parts of this equation

Equation 6 · Part 3 · How Constitutional AI Actually Constrains a Model's Behavior

Symbol E_(x,y_w,y_l)sim D_H cup D_AI

LPM(θ)=− E(x, yw, yl) ∼ DH∪DAI[log⁡σ(rθ(x,yw)−rθ(x,yl))].\mathcal{L}_{PM}(\theta) = -\,\mathbb{E}_{(x,\,y_w,\,y_l)\,\sim\, D_H \cup D_{AI}}\Big[\log \sigma\big(r_\theta(x,y_w) - r_\theta(x,y_l)\big)\Big].
E(x, yw, yl) ∼ DH∪DAI\mathbb{E}_{(x,\,y_w,\,y_l)\,\sim\, D_H \cup D_{AI}}

What this part means

E_(x,ywy_w,yly_l)sim DHD_H cup DAD_AI appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Its job in the formula

E_(x,ywy_w,yly_l)sim DHD_H cup DAD_AI appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

The passage around this formula

Crucially, this AI-generated harmlessness comparison data did not replace human comparison data outright; it was combined with it. The paper reports 135,296 human helpfulness comparisons and 182,831 constitutionally generated harmlessness comparisons feeding one preference model, trained on the union of both, in the authors’ words: “we use human labels for helpfulness, but only AI labels for harmlessness.” A single Bradley-Terry-style preference loss is fit across both sources at once. Writing DHD_H for the human comparison set, DAID_{AI} for the AI-generated set, rθr_\theta for the reward model, x for a prompt, and ywy_w, yly_l for the winning and losing response in a pair, the objective is [displayed…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.