Symbol E
The expected value operator: the probability-weighted average of the quantity inside its brackets.
Read this term in its guide →Published equation contexts
Anthropic’s early red-teaming study gives the most granular public account of this labour market. The team recruited 324 crowdworkers, 307 through Mechanical Turk and 17 through Upwork, and collected 38,961 red-team attacks across model variants, with roughly 11,000 attacks logged against most model types [ 5 ] . Productivity was as concentrated as the labelling data above: about eighty percent of all attacks came from about fifty of the roughly three hundred workers. The paper’s central scaling finding is the one that matters most for this article’s argument — models trained with RLHF became significantly harder to red-team as they grew larger, while plain language models, prompted models,…
The expected value operator: the probability-weighted average of the quantity inside its brackets.
Read this term in its guide →ind is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 18 · AI Safety
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
Anthropic’s early red-teaming study gives the most granular public account of this labour market. The team recruited 324 crowdworkers, 307 through Mechanical Turk and 17 through Upwork, and collected 38,961 red-team attacks across model variants, with roughly 11,000 attacks logged against most model types [ 5 ] . Productivity was as concentrated as the labelling data above: about eighty percent of all attacks came from about fifty of the roughly three hundred workers. The paper’s central scaling finding is the one that matters most for this article’s argument — models trained with RLHF became significantly harder to red-team as they grew larger, while plain language models, prompted models,…