← All parts of this equation

Equation 1 · Part 4 · The Main Technical Approaches to AI Alignment, Compared

Symbol x

P(yA≻yB∣x)=σ(rϕ(x,yA)−rϕ(x,yB)),P(y_A \succ y_B \mid x) = \sigma\big(r_\phi(x,y_A) - r_\phi(x,y_B)\big),
xx

What this part means

x is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Its job in the formula

x is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

The passage around this formula

The procedure fits a scalar reward model to pairwise comparisons using the Bradley–Terry choice model, P(yA≻yB∣x)=σ(rϕ(x,yA)−rϕ(x,yB))P(y_A \succ y_B \mid x) = \sigma\big(r_\phi(x,y_A) - r_\phi(x,y_B)\big). then optimizes the policy against that fitted reward under a penalty that keeps it near its starting point,

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.