← Mathematical compendium

Published equation contexts

PAI(yA≻yB∣x,C)=σ(rψ(x,yA,C)−rψ(x,yB,C))P_{\mathrm{AI}}(y_A \succ y_B \mid x, C) = \sigma\big(r_\psi(x,y_A,C) - r_\psi(x,y_B,C)\big)

Why this formula appears here

Constitutional AI keeps RLHF’s reward-model-and-KL-penalty backbone intact and changes where the comparison labels come from. Bai and colleagues describe a two-phase method: a supervised phase in which the model critiques and revises its own responses against a written set of principles, and a reinforcement phase in which a model, rather than a human, judges which of two candidate responses better satisfies those principles — producing an AI-generated preference dataset that trains the reward model [ 6 ] . Formally, this changes only the source of the comparison label. The Bradley–Terry equation above is unchanged in form; what changes is that the probability being fitted is now [displayed…

Read the full article-specific guide →

Read the representative guide

PAIP_{\mathrm{AI}}

Symbol P_AI

PAP_AI is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →
yAy_A

Symbol y_A

yAy_A is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Read this term in its guide →
yBy_B

Symbol y_B

yBy_B is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

PAI(yA≻yB∣x,C)=σ(rψ(x,yA,C)−rψ(x,yB,C)),P_{\mathrm{AI}}(y_A \succ y_B \mid x, C) = \sigma\big(r_\psi(x,y_A,C) - r_\psi(x,y_B,C)\big),

Equation 5 · AI Safety

The Main Technical Approaches to AI Alignment, Compared

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Constitutional AI keeps RLHF’s reward-model-and-KL-penalty backbone intact and changes where the comparison labels come from. Bai and colleagues describe a two-phase method: a supervised phase in which the model critiques and revises its own responses against a written set of principles, and a reinforcement phase in which a model, rather than a human, judges which of two candidate responses better satisfies those principles — producing an AI-generated preference dataset that trains the reward model [ 6 ] . Formally, this changes only the source of the comparison label. The Bradley–Terry equation above is unchanged in form; what changes is that the probability being fitted is now [displayed…

Equation guide → · Article →