Symbol P_AI
I is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
Constitutional AI keeps RLHF’s reward-model-and-KL-penalty backbone intact and changes where the comparison labels come from. Bai and colleagues describe a two-phase method: a supervised phase in which the model critiques and revises its own responses against a written set of principles, and a reinforcement phase in which a model, rather than a human, judges which of two candidate responses better satisfies those principles — producing an AI-generated preference dataset that trains the reward model [ 6 ] . Formally, this changes only the source of the comparison label. The Bradley–Terry equation above is unchanged in form; what changes is that the probability being fitted is now [displayed…
I is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Read this term in its guide →is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Read this term in its guide →x is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Read this term in its guide →C is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Read this term in its guide →σ is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →si is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 5 · AI Safety
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
Constitutional AI keeps RLHF’s reward-model-and-KL-penalty backbone intact and changes where the comparison labels come from. Bai and colleagues describe a two-phase method: a supervised phase in which the model critiques and revises its own responses against a written set of principles, and a reinforcement phase in which a model, rather than a human, judges which of two candidate responses better satisfies those principles — producing an AI-generated preference dataset that trains the reward model [ 6 ] . Formally, this changes only the source of the comparison label. The Bradley–Terry equation above is unchanged in form; what changes is that the probability being fitted is now [displayed…
Equation guide → · Article →