← Mathematical compendium

Published equation contexts

RBoN(d)=d(αBoN−βBoNd)R_{\mathrm{BoN}}(d) = d(\alpha_{\mathrm{BoN}} - \beta_{\mathrm{BoN}} d)

Why this formula appears here

Gao, Schulman and Hilton measured this cleanly by building a synthetic setting in which a large fixed “gold” reward model plays the role of the human, and a smaller “proxy” reward model is trained on its labels [ 8 ] . Optimising the proxy raises the gold score at first and then lowers it. They fit empirical functional forms in d , the square root of the KL divergence from the initial policy, finding for best-of- n sampling RBoN(d)R_{\mathrm{BoN}}(d) = d(αBoN\alpha_{\mathrm{BoN}} - βBoN\beta_{\mathrm{BoN}} d) and for reinforcement learning

Read the full article-specific guide →

Read the representative guide

RBoNR_{\mathrm{BoN}}

Symbol R_BoN

RBR_BoN is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →
αBoN\alpha_{\mathrm{BoN}}

Symbol alpha_BoN

alphaBa_BoN is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →
βBoN\beta_{\mathrm{BoN}}

Symbol beta_BoN

betaBa_BoN is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

RBoN(d)=d(αBoN−βBoNd)R_{\mathrm{BoN}}(d) = d(\alpha_{\mathrm{BoN}} - \beta_{\mathrm{BoN}} d)

Equation 9 · Alignment & Safety

What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Gao, Schulman and Hilton measured this cleanly by building a synthetic setting in which a large fixed “gold” reward model plays the role of the human, and a smaller “proxy” reward model is trained on its labels [ 8 ] . Optimising the proxy raises the gold score at first and then lowers it. They fit empirical functional forms in d , the square root of the KL divergence from the initial policy, finding for best-of- n sampling RBoN(d)R_{\mathrm{BoN}}(d) = d(αBoN\alpha_{\mathrm{BoN}} - βBoN\beta_{\mathrm{BoN}} d) and for reinforcement learning

Equation guide → · Article →