Symbol R_RL
L is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
Gao, Schulman and Hilton measured this cleanly by building a synthetic setting in which a large fixed “gold” reward model plays the role of the human, and a smaller “proxy” reward model is trained on its labels [ 8 ] . Optimising the proxy raises the gold score at first and then lowers it. They fit empirical functional forms in d , the square root of the KL divergence from the initial policy, finding for best-of- n sampling = d( - d) and for reinforcement learning . with coefficients that vary smoothly and roughly logarithmically with the number of proxy reward model parameters [ 8 ] . Two structural facts are…
L is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →d is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Read this term in its guide →alphL is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →betL is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 10 · Alignment & Safety
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
Gao, Schulman and Hilton measured this cleanly by building a synthetic setting in which a large fixed “gold” reward model plays the role of the human, and a smaller “proxy” reward model is trained on its labels [ 8 ] . Optimising the proxy raises the gold score at first and then lowers it. They fit empirical functional forms in d , the square root of the KL divergence from the initial policy, finding for best-of- n sampling = d( - d) and for reinforcement learning . with coefficients that vary smoothly and roughly logarithmically with the number of proxy reward model parameters [ 8 ] . Two structural facts are…
Equation guide → · Article →