← All parts of this equation

Equation 10 · Part 5 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness

=

RRL(d)=d(αRL−βRLlog⁡d),R_{\mathrm{RL}}(d) = d\left(\alpha_{\mathrm{RL}} - \beta_{\mathrm{RL}} \log d\right),
=

What this part means

The expressions on both sides represent the same quantity under the stated assumptions.

Its job in the formula

The equals sign connects the complete expression on the left with the complete expression on the right. Both sides must have compatible units.

The passage around this formula

Gao, Schulman and Hilton measured this cleanly by building a synthetic setting in which a large fixed “gold” reward model plays the role of the human, and a smaller “proxy” reward model is trained on its labels [ 8 ] . Optimising the proxy raises the gold score at first and then lowers it. They fit empirical functional forms in d , the square root of the KL divergence from the initial policy, finding for best-of- n sampling RBoN(d)R_{\mathrm{BoN}}(d) = d(αBoN\alpha_{\mathrm{BoN}} - βBoN\beta_{\mathrm{BoN}} d) and for reinforcement learning RRL(d)=d(αRL−βRLlog⁡d)R_{\mathrm{RL}}(d) = d\left(\alpha_{\mathrm{RL}} - \beta_{\mathrm{RL}} \log d\right). with coefficients that vary smoothly and roughly logarithmically with the number of proxy reward model parameters [ 8 ] . Two structural facts are…

Read this part in the article →

Learn the underlying idea

An equals sign says that the expression on its left and the expression on its right have the same value under the stated definitions and assumptions.

Open the illustrated equality: what the equals sign claims guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.