Equation 18 · What Alignment Actually Costs
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol E
The expected value operator: the probability-weighted average of the quantity inside its brackets.
Symbol C_find
ind is part of the quantity the equation computes from the expression on the right.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Anthropic’s early red-teaming study gives the most granular public account of this labour market. The team recruited 324 crowdworkers, 307 through Mechanical Turk and 17 through Upwork, and collected 38,961 red-team attacks across model variants, with roughly 11,000 attacks logged against most model types [ 5 ] . Productivity was as concentrated as the labelling data above: about eighty percent of all attacks came from about fifty of the roughly three hundred workers. The paper’s central scaling finding is the one that matters most for this article’s argument — models trained with RLHF became significantly harder to red-team as they grew larger, while plain language models, prompted models,…
Read the full surrounding passage
Anthropic’s early red-teaming study gives the most granular public account of this labour market. The team recruited 324 crowdworkers, 307 through Mechanical Turk and 17 through Upwork, and collected 38,961 red-team attacks across model variants, with roughly 11,000 attacks logged against most model types [ 5 ] . Productivity was as concentrated as the labelling data above: about eighty percent of all attacks came from about fifty of the roughly three hundred workers. The paper’s central scaling finding is the one that matters most for this article’s argument — models trained with RLHF became significantly harder to red-team as they grew larger, while plain language models, prompted models, and rejection-sampling models showed a flat trend with scale [ 5 ] . Put in economic terms, if a single red-team attempt costs c and succeeds with probability p , the expected cost of finding one more genuine failure is . Holding c roughly fixed at crowdworker wage rates, a falling p as RLHF hardens a model means the expected cost per discovered failure rises mechanically with every round of safety training — the same defensive work that red-teaming is meant to validate makes the next round of red-teaming more expensive per useful finding, not less.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to What Alignment Actually Costs