Equation 28 · Why Average Success Rate Hides the Failures That Matter Most
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol q_τ
q_τ is part of the quantity the equation computes from the expression on the right.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
trials in total. This is where the arithmetic becomes unforgiving in a way the mean-difference case never does. Detecting a fixed three-percentage-point difference between two average success rates takes roughly the same, bounded number of trials — on the order of a thousand — regardless of how good either system actually is, because that calculation is about the spread of a difference, not the rarity of an event. Bounding or estimating a rare catastrophic-failure rate is different in kind: the required sample size is inversely proportional to the rate itself. At = 10^{-3} , getting even a rough twenty-event estimate needs on the order of twenty thousand trials. At = 10^{-5} —…
Read the full surrounding passage
trials in total. This is where the arithmetic becomes unforgiving in a way the mean-difference case never does. Detecting a fixed three-percentage-point difference between two average success rates takes roughly the same, bounded number of trials — on the order of a thousand — regardless of how good either system actually is, because that calculation is about the spread of a difference, not the rarity of an event. Bounding or estimating a rare catastrophic-failure rate is different in kind: the required sample size is inversely proportional to the rate itself. At = 10^{-3} , getting even a rough twenty-event estimate needs on the order of twenty thousand trials. At = 10^{-5} — the dangerous-failure probability ceiling regulated safety-critical software is held to under the SIL 4 standard, a bar Rabanser and colleagues cite alongside the FAA’s target of roughly one catastrophic error per billion flight hours as the kind of scale that consequential AI evaluation is implicitly being compared against [ 1 ] — the same twenty-event estimate needs on the order of two million trials. The safer the system claims to be, the more trials it takes to actually confirm that claim at the tail, which is exactly backwards from how confidence usually accrues, and exactly why a modest evaluation budget that comfortably detects an average-case improvement can be wildly insufficient for confirming that a catastrophic failure mode has genuinely been engineered out rather than merely not yet observed.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to Why Average Success Rate Hides the Failures That Matter Most