← Back to article

Equation 26 · Why Average Success Rate Hides the Failures That Matter Most

What does this equation mean?

n≈kqτn \approx \frac{k}{q_\tau}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

nn

Symbol n

n is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

kk

Symbol k

k occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

qτq_\tau

Symbol q_τ

q_τ occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Understand this part →

fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
≈

≈

Approximately equal to; the equality is not exact.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.

What the article says around this equation

​ , so getting a usably precise estimate — a relative error in the neighbourhood of 20 to 25 percent, comparable to observing on the order of twenty events — requires n≈kqτn \approx \frac{k}{q_\tau}. trials in total. This is where the arithmetic becomes unforgiving in a way the mean-difference case never does. Detecting a fixed three-percentage-point difference between two average success rates takes roughly the same, bounded number of trials — on the order of a thousand — regardless of how good either system actually is, because that calculation is about the spread of a difference, not the rarity of an event. Bounding or estimating a rare catastrophic-failure rate is different in kind: the required…
Read the full surrounding passage
​ , so getting a usably precise estimate — a relative error in the neighbourhood of 20 to 25 percent, comparable to observing on the order of twenty events — requires n≈kqτn \approx \frac{k}{q_\tau}. trials in total. This is where the arithmetic becomes unforgiving in a way the mean-difference case never does. Detecting a fixed three-percentage-point difference between two average success rates takes roughly the same, bounded number of trials — on the order of a thousand — regardless of how good either system actually is, because that calculation is about the spread of a difference, not the rarity of an event. Bounding or estimating a rare catastrophic-failure rate is different in kind: the required sample size is inversely proportional to the rate itself. At qτq_\tau = 10^{-3} , getting even a rough twenty-event estimate needs on the order of twenty thousand trials. At qτq_\tau = 10^{-5} — the dangerous-failure probability ceiling regulated safety-critical software is held to under the SIL 4 standard, a bar Rabanser and colleagues cite alongside the FAA’s target of roughly one catastrophic error per billion flight hours as the kind of scale that consequential AI evaluation is implicitly being compared against [ 1 ] — the same twenty-event estimate needs on the order of two million trials. The safer the system claims to be, the more trials it takes to actually confirm that claim at the tail, which is exactly backwards from how confidence usually accrues, and exactly why a modest evaluation budget that comfortably detects an average-case improvement can be wildly insufficient for confirming that a catastrophic failure mode has genuinely been engineered out rather than merely not yet observed.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Why Average Success Rate Hides the Failures That Matter Most

See this formula across 1 published context →

Browse the mathematical compendium →