Equation 21 · What Alignment Actually Costs
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol C_cov
ov is part of the quantity the equation computes from the expression on the right.
Symbol D
D is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Symbol w
w is one factor in the product that computes the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
A useful, explicitly illustrative way to see the size of that gap is to price coverage directly as a scenario, not a reported figure: . the cost of testing D domains to an adequate depth h (in expert-hours per domain) at a specialist wage w . Anchoring h at roughly 150 hours, taken from Anthropic’s single-domain biosecurity effort, and w at a conservative 50 to 100 US dollars an hour for domain expertise — well below typical consulting rates for the specialists these programmes actually recruit — testing even fifty domains to that depth, a small number next to the realistic combination of languages, deployment surfaces, and misuse categories a deployed model faces, comes…
Read the full surrounding passage
A useful, explicitly illustrative way to see the size of that gap is to price coverage directly as a scenario, not a reported figure: . the cost of testing D domains to an adequate depth h (in expert-hours per domain) at a specialist wage w . Anchoring h at roughly 150 hours, taken from Anthropic’s single-domain biosecurity effort, and w at a conservative 50 to 100 US dollars an hour for domain expertise — well below typical consulting rates for the specialists these programmes actually recruit — testing even fifty domains to that depth, a small number next to the realistic combination of languages, deployment surfaces, and misuse categories a deployed model faces, comes to roughly 375,000 to 750,000 US dollars in labour alone for a single testing pass, before any of the iterative retesting that a model updated on any cadence would require. That is a back-of-envelope scenario built from the two anchor figures above, not a cost any lab has published, and it should be read as illustrating the shape of the problem rather than as a documented number: coverage costs scale with the product of domains and depth, both of which the attack surface of a general-purpose model pushes upward faster than any red-teaming budget has grown to match.
Sources cited in the article section
- [5] Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned ↗
- [6] GPT-4 System Card ↗
- [7] Frontier Threats Red Teaming for AI Safety ↗
- [8] Generative Red Team Recap ↗
These citations give research context. Read each source to check which claims it supports.
Return to What Alignment Actually Costs