Equation 4 · Part 1 · Two Different Bets on How to Align a Frontier Model
Symbol p
What this part means
the probability.
Its job in the formula
p is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol p→Article meaning
Where the article explains it
If one crafted probe defeats a given safety-trained policy with probability p , and successive attempts against it were independent, the probability that at least one of k attempts succeeds is = 1 - (1-p)^k, which climbs quickly even when p is small per attempt.
The passage around this formula
which climbs quickly even when p is small per attempt. Two caveats matter as much as the formula. Attempts against one fixed, deployed model are usually correlated rather than independent, so real gains from repeated probing fall below this bound; and a universal, transferable suffix is close to the case the formula flatters…
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the article section
- [9] Universal and Transferable Adversarial Attacks on Aligned Language Models ↗
- [8] Frontier Models are Capable of In-context Scheming ↗
- [10] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training ↗
These citations provide research context; check each source for the exact claim it supports.