Equation 3 · Part 1 · Two Different Bets on How to Align a Frontier Model
Symbol p_k
What this part means
the probability that at least one of k attempts succeeds.
Its job in the formula
is part of the quantity the equation computes from the expression on the right.
Full expression→Symbol p_k→Article meaning
Where the article explains it
If one crafted probe defeats a given safety-trained policy with probability p , and successive attempts against it were independent, the probability that at least one of k attempts succeeds is
The passage around this formula
There is a simple reason a single successful transfer matters more than its raw success rate suggests. If one crafted probe defeats a given safety-trained policy with probability p , and successive attempts against it were independent, the probability that at least one of k attempts succeeds is . which climbs quickly even when p is small per attempt. Two caveats matter as much as the formula. Attempts against one fixed, deployed model are usually correlated rather than independent, so real gains from repeated probing fall below this bound; and a universal, transferable suffix is close to the case the formula flatters most, because it is a single artefact effective across…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the article section
- [9] Universal and Transferable Adversarial Attacks on Aligned Language Models ↗
- [8] Frontier Models are Capable of In-context Scheming ↗
- [10] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training ↗
These citations provide research context; check each source for the exact claim it supports.