Equation 9 · Part 1 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness
Symbol R_BoN
What this part means
oN is part of the quantity the equation computes from the expression on the right.
Its job in the formula
oN is part of the quantity the equation computes from the expression on the right.
Full expression→Symbol R_BoN→Article meaning
The passage around this formula
Gao, Schulman and Hilton measured this cleanly by building a synthetic setting in which a large fixed “gold” reward model plays the role of the human, and a smaller “proxy” reward model is trained on its labels [ 8 ] . Optimising the proxy raises the gold score at first and then lowers it. They fit empirical functional forms in d , the square root of the KL divergence from the initial policy, finding for best-of- n sampling = d( - d) and for reinforcement learning
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.