← Mathematical compendium

Published equation contexts

E[Cfind]=cp\mathbb{E}[C_{\mathrm{find}}] = \frac{c}{p}

Why this formula appears here

Anthropic’s early red-teaming study gives the most granular public account of this labour market. The team recruited 324 crowdworkers, 307 through Mechanical Turk and 17 through Upwork, and collected 38,961 red-team attacks across model variants, with roughly 11,000 attacks logged against most model types [ 5 ] . Productivity was as concentrated as the labelling data above: about eighty percent of all attacks came from about fifty of the roughly three hundred workers. The paper’s central scaling finding is the one that matters most for this article’s argument — models trained with RLHF became significantly harder to red-team as they grew larger, while plain language models, prompted models,…

Read the full article-specific guide →

Read the representative guide

CfindC_{\mathrm{find}}

Symbol C_find

CfC_find is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

E[Cfind]=cp.\mathbb{E}[C_{\mathrm{find}}] = \frac{c}{p}.

Equation 18 · AI Safety

What Alignment Actually Costs

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Anthropic’s early red-teaming study gives the most granular public account of this labour market. The team recruited 324 crowdworkers, 307 through Mechanical Turk and 17 through Upwork, and collected 38,961 red-team attacks across model variants, with roughly 11,000 attacks logged against most model types [ 5 ] . Productivity was as concentrated as the labelling data above: about eighty percent of all attacks came from about fifty of the roughly three hundred workers. The paper’s central scaling finding is the one that matters most for this article’s argument — models trained with RLHF became significantly harder to red-team as they grew larger, while plain language models, prompted models,…

Meanings in this article

  • E\mathbb{E}: The expected value operator: the probability-weighted average of the quantity inside its brackets.
  • cc: the if a single red-team attempt costs.
  • pp: the probability.
Equation guide → · Article →