Equation 12 · Why Average Success Rate Hides the Failures That Matter Most
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol R
R is part of the quantity the equation computes from the expression on the right.
Symbol i
i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Symbol B_i
the blast radius assigned to failure mode i : how much of a system, how much data, how many users or how much liability a given failure could plausibly reach, drawn from a small number of severity tiers rather than treated as a continuous unknown.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
Probability operator
The probability operator gives the chance of the event named inside its brackets or parentheses.
Starting index or lower bound: i
This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Finding the failure is only half the design problem. The other half is scoring it, and a plain pass/fail count throws away exactly the information a consequence-aware evaluation needs — it treats a wrong word choice and a deleted database as the same unit. A more defensible score weights each discovered failure mode by its consequence rather than counting it once: . where is the blast radius assigned to failure mode i : how much of a system, how much data, how many users or how much liability a given failure could plausibly reach, drawn from a small number of severity tiers rather than treated as a continuous unknown. This is not a hypothetical scoring scheme; tiered,…
Read the full surrounding passage
Finding the failure is only half the design problem. The other half is scoring it, and a plain pass/fail count throws away exactly the information a consequence-aware evaluation needs — it treats a wrong word choice and a deleted database as the same unit. A more defensible score weights each discovered failure mode by its consequence rather than counting it once: . where is the blast radius assigned to failure mode i : how much of a system, how much data, how many users or how much liability a given failure could plausibly reach, drawn from a small number of severity tiers rather than treated as a continuous unknown. This is not a hypothetical scoring scheme; tiered, severity-weighted evaluation of exactly this shape is already how the field’s two most cited frontier-risk frameworks operate at the level of an entire model. Anthropic’s Responsible Scaling Policy defines a ladder of AI Safety Levels, modelled loosely on biosafety-level containment standards, under which a model’s required safety, security and deployment safeguards scale with the severity of catastrophic risk its own capability evaluations place it at [ 11 ] . OpenAI’s Preparedness Framework runs a parallel structure at the level of specific risk domains — cybersecurity, chemical/biological/radiological/nuclear capability, persuasion, and model autonomy — scoring each as Low, Medium, High or Critical, and committing not to deploy a model that scores High in a given category until mitigations bring the score back down to Medium [ 12 ] . Neither framework was built for scoring a single agent’s task-level failures; both were built for gating whether an entire model ships. The analogy worth drawing is not that the two problems are identical, but that the field’s own most consequential risk-management structures already reject an averaged, ungraded pass/fail count in favour of exactly this kind of tiered, severity-weighted scoring — and an evaluation harness built for consequential agent tasks has good reason to borrow the same shape at a smaller scale, alongside the general risk-taxonomy work national standards bodies have separately published for generative systems more broadly [ 9 ] .
Sources cited in the surrounding passage
- [11] Anthropic's Responsible Scaling Policy ↗
- [12] Preparedness Framework Version 2 ↗
- [9] Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1) ↗
These citations give research context. Read each source to check which claims it supports.
Return to Why Average Success Rate Hides the Failures That Matter Most