Equation 4 · Two Different Bets on How to Align a Frontier Model
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the probability. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
which climbs quickly even when p is small per attempt. Two caveats matter as much as the formula. Attempts against one fixed, deployed model are usually correlated rather than independent, so real gains from repeated probing fall below this bound; and a universal, transferable suffix is close to the case the formula flatters most, because it is a single artefact effective across many models and many prompts at once rather than a one-off. That asymmetry — a defender must hold every prompt, an attacker only needs one reusable gap — is a structural reason red-teaming keeps finding something regardless of which training mechanism it is pointed at, and it is a poor basis for inferring that the…
Read the full surrounding passage
which climbs quickly even when p is small per attempt. Two caveats matter as much as the formula. Attempts against one fixed, deployed model are usually correlated rather than independent, so real gains from repeated probing fall below this bound; and a universal, transferable suffix is close to the case the formula flatters most, because it is a single artefact effective across many models and many prompts at once rather than a one-off. That asymmetry — a defender must hold every prompt, an attacker only needs one reusable gap — is a structural reason red-teaming keeps finding something regardless of which training mechanism it is pointed at, and it is a poor basis for inferring that the mechanism which was breached first is the weaker one; it may only have been probed first, harder, or by more people.
Sources cited in the article section
- [9] Universal and Transferable Adversarial Attacks on Aligned Language Models ↗
- [8] Frontier Models are Capable of In-context Scheming ↗
- [10] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training ↗
These citations give research context. Read each source to check which claims it supports.
Return to Two Different Bets on How to Align a Frontier Model