Equation 7 · Part 1 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness
Symbol d
What this part means
d is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
d is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol d→Article meaning
The passage around this formula
Gao, Schulman and Hilton measured this cleanly by building a synthetic setting in which a large fixed “gold” reward model plays the role of the human, and a smaller “proxy” reward model is trained on its labels [ 8 ] . Optimising the proxy raises the gold score at first and then lowers it. They fit empirical functional forms in d , the square root of the KL divergence from the initial policy, finding for best-of- n sampling = d( - d) and for reinforcement learning
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.