Equation 1 · Part 6 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness
Symbol r_phi
What this part means
hi is one of the signed contributions combined to compute the quantity on the left.
Its job in the formula
hi is one of the signed contributions combined to compute the quantity on the left.
Full expression→Symbol r_phi→Article meaning
The passage around this formula
A reward model. A network is trained to assign a scalar to a prompt–response pair such that the preferred response scores higher. The near-universal choice is the Bradley–Terry model of paired comparisons, published in Biometrika in 1952 for ranking treatments in incomplete block designs [ 5 ] , which assumes each item has a latent scalar merit and that . with the logistic function. Everything downstream inherits the assumptions in that line. Quality is one-dimensional. Preferences are transitive. Disagreement between raters is noise around a single underlying merit, not evidence of two different merits.
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.