Equation 10 · Part 1 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness
Symbol R_RL
What this part means
L is part of the quantity the equation computes from the expression on the right.
Its job in the formula
L is part of the quantity the equation computes from the expression on the right.
Full expression→Symbol R_RL→Article meaning
The passage around this formula
Gao, Schulman and Hilton measured this cleanly by building a synthetic setting in which a large fixed “gold” reward model plays the role of the human, and a smaller “proxy” reward model is trained on its labels [ 8 ] . Optimising the proxy raises the gold score at first and then lowers it. They fit empirical functional forms in d , the square root of the KL divergence from the initial policy, finding for best-of- n sampling = d( - d) and for reinforcement learning . with coefficients that vary smoothly and roughly logarithmically with the number of proxy reward model parameters [ 8 ] . Two structural facts are…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.