Equation 3 · How AI Security Defenses Actually Work
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol P_hit
it is part of the quantity the equation computes from the expression on the right.
Symbol k
pushed into the thousands, which a machine can do overnight and a human red-team cannot do at all.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The shared structural fact underneath both techniques, and the reason they are described here as one machine rather than two, is what a search process buys once it is automated rather than performed by hand. A single human red-teamer trying prompt variations by hand might realistically test dozens of candidates against a target in a working day. An automated pipeline built on either GCG- or PAIR-style search can attempt thousands of candidates against a running target in the same interval, entirely unattended. If a given search strategy succeeds against a defended target with independent per-attempt probability p , the probability that at least one of k automated attempts succeeds is…
Read the full surrounding passage
The shared structural fact underneath both techniques, and the reason they are described here as one machine rather than two, is what a search process buys once it is automated rather than performed by hand. A single human red-teamer trying prompt variations by hand might realistically test dozens of candidates against a target in a working day. An automated pipeline built on either GCG- or PAIR-style search can attempt thousands of candidates against a running target in the same interval, entirely unattended. If a given search strategy succeeds against a defended target with independent per-attempt probability p , the probability that at least one of k automated attempts succeeds is . Even a strategy with a very low per-attempt success rate becomes a near-certain eventual hit once k is pushed into the thousands, which a machine can do overnight and a human red-team cannot do at all. This is precisely why Anthropic’s red-teaming of its classifier system is measured in thousands of hours rather than a handful of manual sessions [ 5 ] , and it is the direct engineering rationale for building automated red-teaming as a standing pipeline rather than an annual exercise: the defender has to be able to generate and evaluate a search budget that at least approaches what an automated attacker could bring, or the defensive evaluation is measuring the wrong regime entirely.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to How AI Security Defenses Actually Work