Equation 10 · Part 4 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness
Symbol beta_RL
What this part means
betL is one of the signed contributions combined to compute the quantity on the left.
Its job in the formula
betL is one of the signed contributions combined to compute the quantity on the left.
Full expression→Symbol beta_RL→Article meaning
The passage around this formula
Gao, Schulman and Hilton measured this cleanly by building a synthetic setting in which a large fixed “gold” reward model plays the role of the human, and a smaller “proxy” reward model is trained on its labels [ 8 ] . Optimising the proxy raises the gold score at first and then lowers it. They fit empirical functional forms in d , the square root of the KL divergence from the initial policy, finding for best-of- n sampling = d( - d) and for reinforcement learning . with coefficients that vary smoothly and roughly logarithmically with the number of proxy reward model parameters [ 8 ] . Two structural facts are…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.