Equation 10 · Part 2 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness
Symbol d
What this part means
d is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Its job in the formula
d is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Full expression→Symbol d→Article meaning
The passage around this formula
…reward model plays the role of the human, and a smaller “proxy” reward model is trained on its labels [ 8 ] . Optimising the proxy raises the gold score at first and then lowers it. They fit empirical functional forms in d , the square root of the KL divergence from the initial policy, finding for best-of- n sampling = d( - d) and for reinforcement learning . with coefficients that vary smoothly and roughly logarithmically with the…
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.