Equation 2 · How AI Security Defenses Actually Work
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
pushed into the thousands, which a machine can do overnight and a human red-team cannot do at all. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol k
pushed into the thousands, which a machine can do overnight and a human red-team cannot do at all.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The shared structural fact underneath both techniques, and the reason they are described here as one machine rather than two, is what a search process buys once it is automated rather than performed by hand. A single human red-teamer trying prompt variations by hand might realistically test dozens of candidates against a target in a working day. An automated pipeline built on either GCG- or PAIR-style search can attempt thousands of candidates against a running target in the same interval, entirely unattended. If a given search strategy succeeds against a defended target with independent per-attempt probability p , the probability that at least one of k automated attempts succeeds is
Sources cited in the article section
- [6] Red Teaming Language Models with Language Models ↗
- [7] Universal and Transferable Adversarial Attacks on Aligned Language Models ↗
- [8] Jailbreaking Black Box Large Language Models in Twenty Queries ↗
- [5] Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming ↗
These citations give research context. Read each source to check which claims it supports.