← Mathematical compendium

Published equation contexts

RRL(d)=d(αRL−βRLlog⁡d)R_{\mathrm{RL}}(d) = d\left(\alpha_{\mathrm{RL}} - \beta_{\mathrm{RL}} \log d\right)

Why this formula appears here

Gao, Schulman and Hilton measured this cleanly by building a synthetic setting in which a large fixed “gold” reward model plays the role of the human, and a smaller “proxy” reward model is trained on its labels [ 8 ] . Optimising the proxy raises the gold score at first and then lowers it. They fit empirical functional forms in d , the square root of the KL divergence from the initial policy, finding for best-of- n sampling RBoN(d)R_{\mathrm{BoN}}(d) = d(αBoN\alpha_{\mathrm{BoN}} - βBoN\beta_{\mathrm{BoN}} d) and for reinforcement learning RRL(d)=d(αRL−βRLlog⁡d)R_{\mathrm{RL}}(d) = d\left(\alpha_{\mathrm{RL}} - \beta_{\mathrm{RL}} \log d\right). with coefficients that vary smoothly and roughly logarithmically with the number of proxy reward model parameters [ 8 ] . Two structural facts are…

Read the full article-specific guide →

Read the representative guide

RRLR_{\mathrm{RL}}

Symbol R_RL

RRR_RL is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →
αRL\alpha_{\mathrm{RL}}

Symbol alpha_RL

alphaRa_RL is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →
βRL\beta_{\mathrm{RL}}

Symbol beta_RL

betaRa_RL is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

RRL(d)=d(αRL−βRLlog⁡d),R_{\mathrm{RL}}(d) = d\left(\alpha_{\mathrm{RL}} - \beta_{\mathrm{RL}} \log d\right),

Equation 10 · Alignment & Safety

What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Gao, Schulman and Hilton measured this cleanly by building a synthetic setting in which a large fixed “gold” reward model plays the role of the human, and a smaller “proxy” reward model is trained on its labels [ 8 ] . Optimising the proxy raises the gold score at first and then lowers it. They fit empirical functional forms in d , the square root of the KL divergence from the initial policy, finding for best-of- n sampling RBoN(d)R_{\mathrm{BoN}}(d) = d(αBoN\alpha_{\mathrm{BoN}} - βBoN\beta_{\mathrm{BoN}} d) and for reinforcement learning RRL(d)=d(αRL−βRLlog⁡d)R_{\mathrm{RL}}(d) = d\left(\alpha_{\mathrm{RL}} - \beta_{\mathrm{RL}} \log d\right). with coefficients that vary smoothly and roughly logarithmically with the number of proxy reward model parameters [ 8 ] . Two structural facts are…

Equation guide → · Article →