← All parts of this equation

Equation 10 · Part 3 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness

Symbol alpha_RL

RRL(d)=d(αRL−βRLlog⁡d),R_{\mathrm{RL}}(d) = d\left(\alpha_{\mathrm{RL}} - \beta_{\mathrm{RL}} \log d\right),
αRL\alpha_{\mathrm{RL}}

What this part means

alphaRa_RL is one of the signed contributions combined to compute the quantity on the left.

Its job in the formula

alphaRa_RL is one of the signed contributions combined to compute the quantity on the left.

The passage around this formula

Gao, Schulman and Hilton measured this cleanly by building a synthetic setting in which a large fixed “gold” reward model plays the role of the human, and a smaller “proxy” reward model is trained on its labels [ 8 ] . Optimising the proxy raises the gold score at first and then lowers it. They fit empirical functional forms in d , the square root of the KL divergence from the initial policy, finding for best-of- n sampling RBoN(d)R_{\mathrm{BoN}}(d) = d(αBoN\alpha_{\mathrm{BoN}} - βBoN\beta_{\mathrm{BoN}} d) and for reinforcement learning RRL(d)=d(αRL−βRLlog⁡d)R_{\mathrm{RL}}(d) = d\left(\alpha_{\mathrm{RL}} - \beta_{\mathrm{RL}} \log d\right). with coefficients that vary smoothly and roughly logarithmically with the number of proxy reward model parameters [ 8 ] . Two structural facts are…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.