Equation 8 · How Constitutional AI Actually Constrains a Model's Behavior
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the feedback model outputs a probability. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Nothing in that loss distinguishes where a pair came from; a human-labeled winner and an AI-labeled winner are interchangeable once written down as (x, , ) . That is the precise, narrow sense in which Constitutional AI “differs mechanically from plain RLHF”: it changes the labeling function for one half of one dataset, not the loss, not the optimizer, not the use of a KL penalty against the supervised policy. For the soft-labeled AI comparisons specifically, where the feedback model outputs a probability p of preferring response over rather than a hard choice, the corresponding cross-entropy term is
Sources cited in the article section
These citations give research context. Read each source to check which claims it supports.
Return to How Constitutional AI Actually Constrains a Model's Behavior