← Mathematical compendium

Published equation contexts

r(x,y)r(x,y)

Why this formula appears here

The reward term r(x,y) pulls the policy toward whatever the fitted preference model scores highly; the Kullback–Leibler penalty, weighted by β\beta , pulls it back toward the reference distribution. That second term is not a minor regulariser. It is the load-bearing assumption behind the entire method: remove it, and a policy optimised hard enough against an imperfect reward model will find outputs the reward model over-scores without those outputs actually being better, a failure usually called reward hacking. Casper and thirty-one co-authors, surveying RLHF across the field rather than defending any one lab’s implementation, catalogue this and related problems as fundamental rather than…

Read the full article-specific guide →

Read the representative guide

rr

Symbol r

r is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
xx

Symbol x

x is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
yy

Symbol y

y is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (2)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

r(x,y)r(x,y)

Equation 7 · Foundation Models

Claude, From First Principles: Training, Constitutional Methods, and What Actually Shapes a Response

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

The reward term r(x,y) pulls the policy toward whatever the fitted preference model scores highly; the Kullback–Leibler penalty, weighted by β\beta , pulls it back toward the reference distribution. That second term is not a minor regulariser. It is the load-bearing assumption behind the entire method: remove it, and a policy optimised hard enough against an imperfect reward model will find outputs the reward model over-scores without those outputs actually being better, a failure usually called reward hacking. Casper and thirty-one co-authors, surveying RLHF across the field rather than defending any one lab’s implementation, catalogue this and related problems as fundamental rather than…

Equation guide → · Article →
r(x,y)r(x,y)

Equation 10 · Foundation Models

Claude, From First Principles: Training, Constitutional Methods, and What Actually Shapes a Response

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

Constitutional AI is Anthropic’s best-documented departure from plain RLHF, and it is worth being precise about exactly what it changes, because the method is easy either to overstate or to wave at vaguely. Bai and fifty co-authors describe it as a two-phase procedure layered onto the pipeline above [ 1 ] . In the first phase, the model is prompted to critique its own draft response against one principle drawn from a written list, then to revise the response in light of that critique; those revised responses become supervised fine-tuning data, so the model that comes out the other side has been trained partly on its own corrected output rather than solely on external demonstrations. In the…

Equation guide → · Article →