Symbol C_align
lign is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Published equation contexts
This piece assembles the documented cost structure of alignment and safety work at the scale frontier labs now operate, decomposed into three cost centres that map onto three distinct kinds of scarcity: . is the cost of collecting human feedback — demonstrations and preference comparisons, priced per label and bounded by how many people can be recruited and trained. is the compute spent training a policy against a reward model or a verifier, priced in GPU-hours and bounded by the same hardware capability development competes for. is the cost of adversarially testing the result, priced in expert-hours and bounded by…
lign is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →the cost of collecting human feedback — demonstrations and preference comparisons, priced per label and bounded by how many people can be recruited and trained.
Read this term in its guide →the compute spent training a policy against a reward model or a verifier, priced in GPU-hours and bounded by the same hardware capability development competes for.
Read this term in its guide →the cost of adversarially testing the result, priced in expert-hours and bounded by how many domain specialists exist at all.
Read this term in its guide →Its accuracy depends on the assumptions and range of use described in the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 1 · AI Safety
This equation gives an approximation: it relates the quantities while allowing an approximation.
This piece assembles the documented cost structure of alignment and safety work at the scale frontier labs now operate, decomposed into three cost centres that map onto three distinct kinds of scarcity: . is the cost of collecting human feedback — demonstrations and preference comparisons, priced per label and bounded by how many people can be recruited and trained. is the compute spent training a policy against a reward model or a verifier, priced in GPU-hours and bounded by the same hardware capability development competes for. is the cost of adversarially testing the result, priced in expert-hours and bounded by…