← All parts of this equation

Equation 6 · Part 4 · How Constitutional AI Actually Constrains a Model's Behavior

Symbol σ

LPM(θ)=− E(x, yw, yl) ∼ DH∪DAI[log⁡σ(rθ(x,yw)−rθ(x,yl))].\mathcal{L}_{PM}(\theta) = -\,\mathbb{E}_{(x,\,y_w,\,y_l)\,\sim\, D_H \cup D_{AI}}\Big[\log \sigma\big(r_\theta(x,y_w) - r_\theta(x,y_l)\big)\Big].
σ\sigma

What this part means

σ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Its job in the formula

σ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

The passage around this formula

Crucially, this AI-generated harmlessness comparison data did not replace human comparison data outright; it was combined with it. The paper reports 135,296 human helpfulness comparisons and 182,831 constitutionally generated harmlessness comparisons feeding one preference model, trained on the union of both, in the authors’ words: “we use human labels for helpfulness, but only AI labels for harmlessness.” A single Bradley-Terry-style preference loss is fit across both sources at once. Writing DHD_H for the human comparison set, DAID_{AI} for the AI-generated set, rθr_\theta for the reward model, x for a prompt, and ywy_w, yly_l for the winning and losing response in a pair, the objective is [displayed…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.