← Back to article

Equation 17 · What Alignment Actually Costs

What does this equation mean?

pp

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

the probability. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

pp

Symbol p

the probability.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Anthropic’s early red-teaming study gives the most granular public account of this labour market. The team recruited 324 crowdworkers, 307 through Mechanical Turk and 17 through Upwork, and collected 38,961 red-team attacks across model variants, with roughly 11,000 attacks logged against most model types [ 5 ] . Productivity was as concentrated as the labelling data above: about eighty percent of all attacks came from about fifty of the roughly three hundred workers. The paper’s central scaling finding is the one that matters most for this article’s argument — models trained with RLHF became significantly harder to red-team as they grew larger, while plain language models, prompted models,…
Read the full surrounding passage
Anthropic’s early red-teaming study gives the most granular public account of this labour market. The team recruited 324 crowdworkers, 307 through Mechanical Turk and 17 through Upwork, and collected 38,961 red-team attacks across model variants, with roughly 11,000 attacks logged against most model types [ 5 ] . Productivity was as concentrated as the labelling data above: about eighty percent of all attacks came from about fifty of the roughly three hundred workers. The paper’s central scaling finding is the one that matters most for this article’s argument — models trained with RLHF became significantly harder to red-team as they grew larger, while plain language models, prompted models, and rejection-sampling models showed a flat trend with scale [ 5 ] . Put in economic terms, if a single red-team attempt costs c and succeeds with probability p , the expected cost of finding one more genuine failure is

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to What Alignment Actually Costs

Browse the mathematical compendium →