← All parts of this equation

Equation 9 · Part 2 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness

Symbol d

RBoN(d)=d(αBoN−βBoNd)R_{\mathrm{BoN}}(d) = d(\alpha_{\mathrm{BoN}} - \beta_{\mathrm{BoN}} d)
dd

What this part means

d is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Its job in the formula

d is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

The passage around this formula

Gao, Schulman and Hilton measured this cleanly by building a synthetic setting in which a large fixed “gold” reward model plays the role of the human, and a smaller “proxy” reward model is trained on its labels [ 8 ] . Optimising the proxy raises the gold score at first and then lowers it. They fit empirical functional forms in d , the square root of the KL divergence from the initial policy, finding for best-of- n sampling RBoN(d)R_{\mathrm{BoN}}(d) = d(αBoN\alpha_{\mathrm{BoN}} - βBoN\beta_{\mathrm{BoN}} d) and for reinforcement learning

Read this part in the article →

Learn the underlying idea

A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.

Open the illustrated functions: inputs become outputs guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.