← Back to article

Equation 3 · Two Different Bets on How to Align a Frontier Model

What does this equation mean?

pk=1−(1−p)k,p_k = 1 - (1-p)^k,

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operations1 - (1-p)^k
Result or conditionp_k
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

pkp_k

Symbol p_k

the probability that at least one of k attempts succeeds.

Understand this part →

pp

Symbol p

the probability.

Understand this part →

kk

Symbol k

k is part of the quantity the equation computes from the expression on the right.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
subtraction

subtraction

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

There is a simple reason a single successful transfer matters more than its raw success rate suggests. If one crafted probe defeats a given safety-trained policy with probability p , and successive attempts against it were independent, the probability that at least one of k attempts succeeds is pk=1−(1−p)kp_k = 1 - (1-p)^k. which climbs quickly even when p is small per attempt. Two caveats matter as much as the formula. Attempts against one fixed, deployed model are usually correlated rather than independent, so real gains from repeated probing fall below this bound; and a universal, transferable suffix is close to the case the formula flatters most, because it is a single artefact effective across…
Read the full surrounding passage
There is a simple reason a single successful transfer matters more than its raw success rate suggests. If one crafted probe defeats a given safety-trained policy with probability p , and successive attempts against it were independent, the probability that at least one of k attempts succeeds is pk=1−(1−p)kp_k = 1 - (1-p)^k. which climbs quickly even when p is small per attempt. Two caveats matter as much as the formula. Attempts against one fixed, deployed model are usually correlated rather than independent, so real gains from repeated probing fall below this bound; and a universal, transferable suffix is close to the case the formula flatters most, because it is a single artefact effective across many models and many prompts at once rather than a one-off. That asymmetry — a defender must hold every prompt, an attacker only needs one reusable gap — is a structural reason red-teaming keeps finding something regardless of which training mechanism it is pointed at, and it is a poor basis for inferring that the mechanism which was breached first is the weaker one; it may only have been probed first, harder, or by more people.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to Two Different Bets on How to Align a Frontier Model

See this formula across 3 published contexts →

Browse the mathematical compendium →