← Mathematical compendium

Published equation contexts

P(yA preferred to yB∣x)=σ ⁣(rϕ(x,yA)−rϕ(x,yB))P(y_A \text{ preferred to } y_B \mid x) = \sigma\!\left(r_\phi(x, y_A) - r_\phi(x, y_B)\right)

Why this formula appears here

A reward model. A network is trained to assign a scalar to a prompt–response pair such that the preferred response scores higher. The near-universal choice is the Bradley–Terry model of paired comparisons, published in Biometrika in 1952 for ranking treatments in incomplete block designs [ 5 ] , which assumes each item has a latent scalar merit and that P(yA preferred to yB∣x)=σ ⁣(rϕ(x,yA)−rϕ(x,yB))P(y_A \text{ preferred to } y_B \mid x) = \sigma\!\left(r_\phi(x, y_A) - r_\phi(x, y_B)\right). with σ\sigma the logistic function. Everything downstream inherits the assumptions in that line. Quality is one-dimensional. Preferences are transitive. Disagreement between raters is noise around a single underlying merit, not evidence of two different merits.

Read the full article-specific guide →

Read the representative guide

yAy_A

Symbol y_A

yAy_A is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Read this term in its guide →
yBy_B

Symbol y_B

yBy_B is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

P(yA preferred to yB∣x)=σ ⁣(rϕ(x,yA)−rϕ(x,yB)),P(y_A \text{ preferred to } y_B \mid x) = \sigma\!\left(r_\phi(x, y_A) - r_\phi(x, y_B)\right),

Equation 1 · Alignment & Safety

What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

A reward model. A network is trained to assign a scalar to a prompt–response pair such that the preferred response scores higher. The near-universal choice is the Bradley–Terry model of paired comparisons, published in Biometrika in 1952 for ranking treatments in incomplete block designs [ 5 ] , which assumes each item has a latent scalar merit and that P(yA preferred to yB∣x)=σ ⁣(rϕ(x,yA)−rϕ(x,yB))P(y_A \text{ preferred to } y_B \mid x) = \sigma\!\left(r_\phi(x, y_A) - r_\phi(x, y_B)\right). with σ\sigma the logistic function. Everything downstream inherits the assumptions in that line. Quality is one-dimensional. Preferences are transitive. Disagreement between raters is noise around a single underlying merit, not evidence of two different merits.

Meanings in this article

Equation guide → · Article →