← Back to article

Equation 4 · How AI Security Defenses Actually Work

What does this equation mean?

kk

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

pushed into the thousands, which a machine can do overnight and a human red-team cannot do at all. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

kk

Symbol k

pushed into the thousands, which a machine can do overnight and a human red-team cannot do at all.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Even a strategy with a very low per-attempt success rate becomes a near-certain eventual hit once k is pushed into the thousands, which a machine can do overnight and a human red-team cannot do at all. This is precisely why Anthropic’s red-teaming of its classifier system is measured in thousands of hours rather than a handful of manual sessions [ 5 ] , and it is the direct engineering rationale for building automated red-teaming as a standing pipeline rather than an annual exercise: the defender has to be able to generate and evaluate a search budget that at least approaches what an automated attacker could bring, or the defensive evaluation is measuring the wrong regime entirely.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to How AI Security Defenses Actually Work

Browse the mathematical compendium →