Equation 6 · Part 1 · How Constitutional AI Actually Constrains a Model's Behavior
Symbol L_PM
What this part means
M is computed from the expected values combined on the right.
Its job in the formula
M is computed from the expected values combined on the right.
Full expression→Symbol L_PM→Article meaning
The passage around this formula
Crucially, this AI-generated harmlessness comparison data did not replace human comparison data outright; it was combined with it. The paper reports 135,296 human helpfulness comparisons and 182,831 constitutionally generated harmlessness comparisons feeding one preference model, trained on the union of both, in the authors’ words: “we use human labels for helpfulness, but only AI labels for harmlessness.” A single Bradley-Terry-style preference loss is fit across both sources at once. Writing for the human comparison set, for the AI-generated set, for the reward model, x for a prompt, and , for the winning and losing response in a pair, the objective is [displayed…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the article section
These citations provide research context; check each source for the exact claim it supports.