Equation 5 · Part 2 · How Constitutional AI Actually Constrains a Model's Behavior
Symbol y_l
What this part means
is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol y_l→Article meaning
The passage around this formula
Crucially, this AI-generated harmlessness comparison data did not replace human comparison data outright; it was combined with it. The paper reports 135,296 human helpfulness comparisons and 182,831 constitutionally generated harmlessness comparisons feeding one preference model, trained on the union of both, in the authors’ words: “we use human labels for helpfulness, but only AI labels for harmlessness.” A single Bradley-Terry-style preference loss is fit across both sources at once. Writing for the human comparison set, for the AI-generated set, for the reward model, x for a prompt, and , for the winning and losing response in a pair, the objective is
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the article section
These citations provide research context; check each source for the exact claim it supports.