← All parts of this equation

Equation 1 · Part 3 · The Main Technical Approaches to AI Alignment, Compared

Symbol y_B

P(yA≻yB∣x)=σ(rϕ(x,yA)−rϕ(x,yB)),P(y_A \succ y_B \mid x) = \sigma\big(r_\phi(x,y_A) - r_\phi(x,y_B)\big),
yBy_B

What this part means

yBy_B is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Its job in the formula

yBy_B is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

The passage around this formula

The procedure fits a scalar reward model to pairwise comparisons using the Bradley–Terry choice model, P(yA≻yB∣x)=σ(rϕ(x,yA)−rϕ(x,yB))P(y_A \succ y_B \mid x) = \sigma\big(r_\phi(x,y_A) - r_\phi(x,y_B)\big). then optimizes the policy against that fitted reward under a penalty that keeps it near its starting point,

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.