← Back to article

Equation 1 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness

What does this equation mean?

P(yA preferred to yB∣x)=σ ⁣(rϕ(x,yA)−rϕ(x,yB)),P(y_A \text{ preferred to } y_B \mid x) = \sigma\!\left(r_\phi(x, y_A) - r_\phi(x, y_B)\right),

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationsσ(r_phi(x, y_A) - r_phi(x, y_B))
Result or conditionP(y_A preferred to y_B mid x)
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

PP

Symbol P

P is part of the quantity the equation computes from the expression on the right.

Understand this part →

yAy_A

Symbol y_A

yAy_A is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

yBy_B

Symbol y_B

yBy_B is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

xx

Symbol x

x is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

σ\sigma

Symbol σ

the logistic function.

Understand this part →

rϕr_\phi

Symbol r_phi

rpr_phi is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
subtraction

subtraction

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

A reward model. A network is trained to assign a scalar to a prompt–response pair such that the preferred response scores higher. The near-universal choice is the Bradley–Terry model of paired comparisons, published in Biometrika in 1952 for ranking treatments in incomplete block designs [ 5 ] , which assumes each item has a latent scalar merit and that P(yA preferred to yB∣x)=σ ⁣(rϕ(x,yA)−rϕ(x,yB))P(y_A \text{ preferred to } y_B \mid x) = \sigma\!\left(r_\phi(x, y_A) - r_\phi(x, y_B)\right). with σ\sigma the logistic function. Everything downstream inherits the assumptions in that line. Quality is one-dimensional. Preferences are transitive. Disagreement between raters is noise around a single underlying merit, not evidence of two different merits.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness

See this formula across 1 published context →

Browse the mathematical compendium →