Symbol θ
θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Published equation contexts
then optimizes the policy against that fitted reward under a penalty that keeps it near its starting point, . Read the two equations together and the oversight limit is visible in the structure itself, not only in later critiques of it. The reward model is trained on whatever a human rater could see and judge in the time given; nothing forces it to track truth or usefulness beyond what raters could distinguish. The KL term in the second equation is a standing admission of that limit — it constrains how far the policy may travel from the region where was actually calibrated, because past that region the fitted proxy is not trusted to mean anything. Casper and…
θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →, ysimpi_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →hi is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →x is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →y is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →β is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →L is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →pi_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →pef is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Read this expression with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 2 · AI Safety
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
then optimizes the policy against that fitted reward under a penalty that keeps it near its starting point, . Read the two equations together and the oversight limit is visible in the structure itself, not only in later critiques of it. The reward model is trained on whatever a human rater could see and judge in the time given; nothing forces it to track truth or usefulness beyond what raters could distinguish. The KL term in the second equation is a standing admission of that limit — it constrains how far the policy may travel from the region where was actually calibrated, because past that region the fitted proxy is not trusted to mean anything. Casper and…
Equation guide → · Article →