← Back to article

Equation 1 · What Alignment Actually Costs

What does this equation mean?

Calign≈Chuman+CRL+CredteamC_{\mathrm{align}} \approx C_{\mathrm{human}} + C_{\mathrm{RL}} + C_{\mathrm{redteam}}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

CalignC_{\mathrm{align}}

Symbol C_align

CaC_align is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

ChumanC_{\mathrm{human}}

Symbol C_human

the cost of collecting human feedback — demonstrations and preference comparisons, priced per label and bounded by how many people can be recruited and trained.

Understand this part →

CRLC_{\mathrm{RL}}

Symbol C_RL

the compute spent training a policy against a reward model or a verifier, priced in GPU-hours and bounded by the same hardware capability development competes for.

Understand this part →

CredteamC_{\mathrm{redteam}}

Symbol C_redteam

the cost of adversarially testing the result, priced in expert-hours and bounded by how many domain specialists exist at all.

Understand this part →

≈

≈

Approximately equal to; the equality is not exact.

Understand this part →

addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Its accuracy depends on the assumptions and range of use described in the article.

What the article says around this equation

This piece assembles the documented cost structure of alignment and safety work at the scale frontier labs now operate, decomposed into three cost centres that map onto three distinct kinds of scarcity: Calign≈Chuman+CRL+CredteamC_{\mathrm{align}} \approx C_{\mathrm{human}} + C_{\mathrm{RL}} + C_{\mathrm{redteam}}. ChumanC_{\mathrm{human}} is the cost of collecting human feedback — demonstrations and preference comparisons, priced per label and bounded by how many people can be recruited and trained. CRLC_{\mathrm{RL}} is the compute spent training a policy against a reward model or a verifier, priced in GPU-hours and bounded by the same hardware capability development competes for. CredteamC_{\mathrm{redteam}} is the cost of adversarially testing the result, priced in expert-hours and bounded by…
Read the full surrounding passage
This piece assembles the documented cost structure of alignment and safety work at the scale frontier labs now operate, decomposed into three cost centres that map onto three distinct kinds of scarcity: Calign≈Chuman+CRL+CredteamC_{\mathrm{align}} \approx C_{\mathrm{human}} + C_{\mathrm{RL}} + C_{\mathrm{redteam}}. ChumanC_{\mathrm{human}} is the cost of collecting human feedback — demonstrations and preference comparisons, priced per label and bounded by how many people can be recruited and trained. CRLC_{\mathrm{RL}} is the compute spent training a policy against a reward model or a verifier, priced in GPU-hours and bounded by the same hardware capability development competes for. CredteamC_{\mathrm{redteam}} is the cost of adversarially testing the result, priced in expert-hours and bounded by how many domain specialists exist at all. Each term has published figures behind it. None of them is small in absolute terms. All of them are small relative to what they are meant to check.

Read the equation in its article →

For background, read the article’s source list.

Return to What Alignment Actually Costs

See this formula across 1 published context →

Browse the mathematical compendium →