← All parts of this equation

Equation 3 · Part 1 · How Constitutional AI Actually Constrains a Model's Behavior

Symbol r_θ

rθr_\theta
rθr_\theta

What this part means

r_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Its job in the formula

r_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

The passage around this formula

Crucially, this AI-generated harmlessness comparison data did not replace human comparison data outright; it was combined with it. The paper reports 135,296 human helpfulness comparisons and 182,831 constitutionally generated harmlessness comparisons feeding one preference model, trained on the union of both, in the authors’ words: “we use human labels for helpfulness, but only AI labels for harmlessness.” A single Bradley-Terry-style preference loss is fit across both sources at once. Writing DHD_H for the human comparison set, DAID_{AI} for the AI-generated set, rθr_\theta for the reward model, x for a prompt, and ywy_w, yly_l for the winning and losing response in a pair, the objective is

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.