Symbol L_PM
M is computed from the expected values combined on the right.
Read this term in its guide →Published equation contexts
Crucially, this AI-generated harmlessness comparison data did not replace human comparison data outright; it was combined with it. The paper reports 135,296 human helpfulness comparisons and 182,831 constitutionally generated harmlessness comparisons feeding one preference model, trained on the union of both, in the authors’ words: “we use human labels for helpfulness, but only AI labels for harmlessness.” A single Bradley-Terry-style preference loss is fit across both sources at once. Writing for the human comparison set, for the AI-generated set, for the reward model, x for a prompt, and , for the winning and losing response in a pair, the objective is [displayed…
M is computed from the expected values combined on the right.
Read this term in its guide →θ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Read this term in its guide →E_(x,,)sim cup I appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →σ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →r_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →x appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 6 · Foundation Models
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
Crucially, this AI-generated harmlessness comparison data did not replace human comparison data outright; it was combined with it. The paper reports 135,296 human helpfulness comparisons and 182,831 constitutionally generated harmlessness comparisons feeding one preference model, trained on the union of both, in the authors’ words: “we use human labels for helpfulness, but only AI labels for harmlessness.” A single Bradley-Terry-style preference loss is fit across both sources at once. Writing for the human comparison set, for the AI-generated set, for the reward model, x for a prompt, and , for the winning and losing response in a pair, the objective is [displayed…
Equation guide → · Article →