← All parts of this equation

Equation 6 · Part 5 · How Constitutional AI Actually Constrains a Model's Behavior

Symbol r_θ

LPM(θ)=− E(x, yw, yl) ∼ DH∪DAI[log⁡σ(rθ(x,yw)−rθ(x,yl))].\mathcal{L}_{PM}(\theta) = -\,\mathbb{E}_{(x,\,y_w,\,y_l)\,\sim\, D_H \cup D_{AI}}\Big[\log \sigma\big(r_\theta(x,y_w) - r_\theta(x,y_l)\big)\Big].
rθr_\theta

What this part means

r_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Its job in the formula

r_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

The passage around this formula

…labels for helpfulness, but only AI labels for harmlessness.” A single Bradley-Terry-style preference loss is fit across both sources at once. Writing DHD_H for the human comparison set, DAID_{AI} for the AI-generated set, rθr_\theta for the reward model, x for a prompt, and ywy_w, yly_l for the winning and losing response in a pair, the objective is LPM(θ)=− E(x, yw, yl) ∼ DH∪DAI[log⁡σ(rθ(x,yw)−rθ(x,yl))]\mathcal{L}_{PM}(\theta) = -\,\mathbb{E}_{(x,\,y_w,\,y_l)\,\sim\, D_H \cup D_{AI}}\Big[\log \sigma\big(r_\theta(x,y_w) - r_\theta(x,y_l)\big)\Big]. Nothing in that loss distinguishes where a pair came from; a human-labeled winner and an AI-labeled winner are interchangeable once written down as (x,…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.