Symbol p
its failure probability, and let severity be a random quantity C 0 describing how bad a given failure turns out to be, conditional on failure occurring.
Read this term in its guide →Published equation contexts
It helps to state the structural point precisely rather than just by example. Let p be an agent’s probability of completing a task successfully, so 1-p is its failure probability, and let severity be a random quantity C 0 describing how bad a given failure turns out to be, conditional on failure occurring. The probability that a given run produces a failure whose severity clears some consequential threshold — a deleted database rather than a clumsy sentence — is
its failure probability, and let severity be a random quantity C 0 describing how bad a given failure turns out to be, conditional on failure occurring.
Read this term in its guide →Read this expression with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 2 · Model Evaluation
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
It helps to state the structural point precisely rather than just by example. Let p be an agent’s probability of completing a task successfully, so 1-p is its failure probability, and let severity be a random quantity C 0 describing how bad a given failure turns out to be, conditional on failure occurring. The probability that a given run produces a failure whose severity clears some consequential threshold — a deleted database rather than a clumsy sentence — is
Equation 1 · AI Safety
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
The failure modes built on repeated querying — jailbreak search, paraphrase evasion, mismatched-language generalisation — share a further structural property worth stating precisely, because a single strong per-attempt success rate is often reported as though it settled the question of robustness on its own. Suppose a safeguard blocks a fixed, independently generated attack attempt with probability 1-p , so a single try succeeds only with probability p . An adversary who is not limited to one try, and who runs k independent attempts — exactly the loop PAIR and the suffix-search method both implement mechanically [ 2 , 1 ] — succeeds at least once with probability