Equation 6 · Part 2 · Claude, From First Principles: Training, Constitutional Methods, and What Actually Shapes a Response
Symbol θ
What this part means
θ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Its job in the formula
θ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Full expression→Symbol θ→Article meaning
The passage around this formula
The optimisation itself has a specific and revealing shape. Rather than maximising the fitted reward outright, the standard formulation constrains the updated policy to stay close to a reference policy, ordinarily the model just after supervised fine-tuning, one step removed from the raw pretrained prior: . The reward term r(x,y) pulls the policy toward whatever the fitted preference model scores highly; the Kullback–Leibler penalty, weighted by , pulls it back toward the reference distribution. That second term is not a minor regulariser. It is the load-bearing assumption behind the entire method: remove it, and a policy optimised hard enough against an imperfect…
Learn the underlying idea
A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.
Open the illustrated functions: inputs become outputs guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.