Equation 1 · What Alignment Actually Costs
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol C_align
lign is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol C_human
the cost of collecting human feedback — demonstrations and preference comparisons, priced per label and bounded by how many people can be recruited and trained.
Symbol C_RL
the compute spent training a policy against a reward model or a verifier, priced in GPU-hours and bounded by the same hardware capability development competes for.
Symbol C_redteam
the cost of adversarially testing the result, priced in expert-hours and bounded by how many domain specialists exist at all.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Its accuracy depends on the assumptions and range of use described in the article.
What the article says around this equation
This piece assembles the documented cost structure of alignment and safety work at the scale frontier labs now operate, decomposed into three cost centres that map onto three distinct kinds of scarcity: . is the cost of collecting human feedback — demonstrations and preference comparisons, priced per label and bounded by how many people can be recruited and trained. is the compute spent training a policy against a reward model or a verifier, priced in GPU-hours and bounded by the same hardware capability development competes for. is the cost of adversarially testing the result, priced in expert-hours and bounded by…
Read the full surrounding passage
This piece assembles the documented cost structure of alignment and safety work at the scale frontier labs now operate, decomposed into three cost centres that map onto three distinct kinds of scarcity: . is the cost of collecting human feedback — demonstrations and preference comparisons, priced per label and bounded by how many people can be recruited and trained. is the compute spent training a policy against a reward model or a verifier, priced in GPU-hours and bounded by the same hardware capability development competes for. is the cost of adversarially testing the result, priced in expert-hours and bounded by how many domain specialists exist at all. Each term has published figures behind it. None of them is small in absolute terms. All of them are small relative to what they are meant to check.
For background, read the article’s source list.
Return to What Alignment Actually Costs