Equation 10 · Claude, From First Principles: Training, Constitutional Methods, and What Actually Shapes a Response
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol r
r is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol x
x is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol y
y is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Constitutional AI is Anthropic’s best-documented departure from plain RLHF, and it is worth being precise about exactly what it changes, because the method is easy either to overstate or to wave at vaguely. Bai and fifty co-authors describe it as a two-phase procedure layered onto the pipeline above [ 1 ] . In the first phase, the model is prompted to critique its own draft response against one principle drawn from a written list, then to revise the response in light of that critique; those revised responses become supervised fine-tuning data, so the model that comes out the other side has been trained partly on its own corrected output rather than solely on external demonstrations. In the…
Read the full surrounding passage
Constitutional AI is Anthropic’s best-documented departure from plain RLHF, and it is worth being precise about exactly what it changes, because the method is easy either to overstate or to wave at vaguely. Bai and fifty co-authors describe it as a two-phase procedure layered onto the pipeline above [ 1 ] . In the first phase, the model is prompted to critique its own draft response against one principle drawn from a written list, then to revise the response in light of that critique; those revised responses become supervised fine-tuning data, so the model that comes out the other side has been trained partly on its own corrected output rather than solely on external demonstrations. In the second phase, the same reinforcement-learning-against-a-reward-model structure from the previous section is reused, but for the harmlessness comparisons specifically, the preference labels ordinarily supplied by human raters are instead generated by a separate instance of the model, prompted with a constitutional principle and asked which of two candidate responses better satisfies it — a substitution the paper’s authors call reinforcement learning from AI feedback, RLAIF, to distinguish it from RLHF [ 1 ] . Put in terms of the objective in the previous section, Constitutional AI’s second phase is the same J() , with r(x,y) for the harmlessness component sourced from a principle-conditioned AI judgment rather than an additional round of human labelling; the original paper’s helpfulness comparisons continued to draw on human preference data [ 1 ] .
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.