← All parts of this equation

Equation 1 · Part 7 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness

=

P(yA preferred to yB∣x)=σ ⁣(rϕ(x,yA)−rϕ(x,yB)),P(y_A \text{ preferred to } y_B \mid x) = \sigma\!\left(r_\phi(x, y_A) - r_\phi(x, y_B)\right),
=

What this part means

The expressions on both sides represent the same quantity under the stated assumptions.

Its job in the formula

The equals sign connects the complete expression on the left with the complete expression on the right. Both sides must have compatible units.

The passage around this formula

A reward model. A network is trained to assign a scalar to a prompt–response pair such that the preferred response scores higher. The near-universal choice is the Bradley–Terry model of paired comparisons, published in Biometrika in 1952 for ranking treatments in incomplete block designs [ 5 ] , which assumes each item has a latent scalar merit and that P(yA preferred to yB∣x)=σ ⁣(rϕ(x,yA)−rϕ(x,yB))P(y_A \text{ preferred to } y_B \mid x) = \sigma\!\left(r_\phi(x, y_A) - r_\phi(x, y_B)\right). with σ\sigma the logistic function. Everything downstream inherits the assumptions in that line. Quality is one-dimensional. Preferences are transitive. Disagreement between raters is noise around a single underlying merit, not evidence of two different merits.

Read this part in the article →

Learn the underlying idea

An equals sign says that the expression on its left and the expression on its right have the same value under the stated definitions and assumptions.

Open the illustrated equality: what the equals sign claims guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.