Equation 4 · Ten Failure Modes That Define Deployed AI Safety Systems
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol P
P is part of the quantity the equation computes from the expression on the right.
Symbol k
k is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The failure modes built on repeated querying — jailbreak search, paraphrase evasion, mismatched-language generalisation — share a further structural property worth stating precisely, because a single strong per-attempt success rate is often reported as though it settled the question of robustness on its own. Suppose a safeguard blocks a fixed, independently generated attack attempt with probability 1-p , so a single try succeeds only with probability p . An adversary who is not limited to one try, and who runs k independent attempts — exactly the loop PAIR and the suffix-search method both implement mechanically [ 2 , 1 ] — succeeds at least once with probability . At p =…
Read the full surrounding passage
The failure modes built on repeated querying — jailbreak search, paraphrase evasion, mismatched-language generalisation — share a further structural property worth stating precisely, because a single strong per-attempt success rate is often reported as though it settled the question of robustness on its own. Suppose a safeguard blocks a fixed, independently generated attack attempt with probability 1-p , so a single try succeeds only with probability p . An adversary who is not limited to one try, and who runs k independent attempts — exactly the loop PAIR and the suffix-search method both implement mechanically [ 2 , 1 ] — succeeds at least once with probability . At p = 0.05 , a safeguard that looks 95% reliable per attempt, twenty tries already succeed with probability 1-(0.95)^{20} 0.64 , and by sixty tries cumulative success exceeds 0.95 — a per-attempt figure that reads as strong evidence of robustness collapses once retrying is free, which is precisely the black-box search loop several of the documented attacks above run automatically. This is not a result reported by any single cited study; it is the elementary shape of what those studies demonstrate empirically, and it is the reason a reported single-attempt block rate, on its own, answers a narrower question than it is usually taken to answer.
For background, read the article’s source list.
Return to Ten Failure Modes That Define Deployed AI Safety Systems