Equation 25 · What Alignment Actually Costs
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the anchoring. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
the cost of testing D domains to an adequate depth h (in expert-hours per domain) at a specialist wage w . Anchoring h at roughly 150 hours, taken from Anthropic’s single-domain biosecurity effort, and w at a conservative 50 to 100 US dollars an hour for domain expertise — well below typical consulting rates for the specialists these programmes actually recruit — testing even fifty domains to that depth, a small number next to the realistic combination of languages, deployment surfaces, and misuse categories a deployed model faces, comes to roughly 375,000 to 750,000 US dollars in labour alone for a single testing pass, before any of the iterative retesting that a model updated on any…
Read the full surrounding passage
the cost of testing D domains to an adequate depth h (in expert-hours per domain) at a specialist wage w . Anchoring h at roughly 150 hours, taken from Anthropic’s single-domain biosecurity effort, and w at a conservative 50 to 100 US dollars an hour for domain expertise — well below typical consulting rates for the specialists these programmes actually recruit — testing even fifty domains to that depth, a small number next to the realistic combination of languages, deployment surfaces, and misuse categories a deployed model faces, comes to roughly 375,000 to 750,000 US dollars in labour alone for a single testing pass, before any of the iterative retesting that a model updated on any cadence would require. That is a back-of-envelope scenario built from the two anchor figures above, not a cost any lab has published, and it should be read as illustrating the shape of the problem rather than as a documented number: coverage costs scale with the product of domains and depth, both of which the attack surface of a general-purpose model pushes upward faster than any red-teaming budget has grown to match.
Sources cited in the article section
- [5] Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned ↗
- [6] GPT-4 System Card ↗
- [7] Frontier Threats Red Teaming for AI Safety ↗
- [8] Generative Red Team Recap ↗
These citations give research context. Read each source to check which claims it supports.