Equation 26 · Why Average Success Rate Hides the Failures That Matter Most
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol n
n is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol k
k occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Symbol q_τ
q_τ occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.
What the article says around this equation
, so getting a usably precise estimate — a relative error in the neighbourhood of 20 to 25 percent, comparable to observing on the order of twenty events — requires . trials in total. This is where the arithmetic becomes unforgiving in a way the mean-difference case never does. Detecting a fixed three-percentage-point difference between two average success rates takes roughly the same, bounded number of trials — on the order of a thousand — regardless of how good either system actually is, because that calculation is about the spread of a difference, not the rarity of an event. Bounding or estimating a rare catastrophic-failure rate is different in kind: the required…
Read the full surrounding passage
, so getting a usably precise estimate — a relative error in the neighbourhood of 20 to 25 percent, comparable to observing on the order of twenty events — requires . trials in total. This is where the arithmetic becomes unforgiving in a way the mean-difference case never does. Detecting a fixed three-percentage-point difference between two average success rates takes roughly the same, bounded number of trials — on the order of a thousand — regardless of how good either system actually is, because that calculation is about the spread of a difference, not the rarity of an event. Bounding or estimating a rare catastrophic-failure rate is different in kind: the required sample size is inversely proportional to the rate itself. At = 10^{-3} , getting even a rough twenty-event estimate needs on the order of twenty thousand trials. At = 10^{-5} — the dangerous-failure probability ceiling regulated safety-critical software is held to under the SIL 4 standard, a bar Rabanser and colleagues cite alongside the FAA’s target of roughly one catastrophic error per billion flight hours as the kind of scale that consequential AI evaluation is implicitly being compared against [ 1 ] — the same twenty-event estimate needs on the order of two million trials. The safer the system claims to be, the more trials it takes to actually confirm that claim at the tail, which is exactly backwards from how confidence usually accrues, and exactly why a modest evaluation budget that comfortably detects an average-case improvement can be wildly insufficient for confirming that a catastrophic failure mode has genuinely been engineered out rather than merely not yet observed.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to Why Average Success Rate Hides the Failures That Matter Most