Equation 6 · How Do We Actually Know a Safety Measure Worked?
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol N
N is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
A more accountable claim would state the access budget testers were given, not only the hours logged — prototypes or public variants, live scores or a black box, architecture documentation or none, because the Constitutional Classifiers case shows this dimension, not effort alone, was what separated the two results. It would name which of the four red-teaming quadrants — human or automated, structured or unstructured — actually ran, since Perez’s automated volume, Ganguli’s human panels, and AISI’s threat-specific structured teams each found different things, and none of the three would have found what the others did. It would distinguish, in the same sentence, a production side-effect…
Read the full surrounding passage
A more accountable claim would state the access budget testers were given, not only the hours logged — prototypes or public variants, live scores or a black box, architecture documentation or none, because the Constitutional Classifiers case shows this dimension, not effort alone, was what separated the two results. It would name which of the four red-teaming quadrants — human or automated, structured or unstructured — actually ran, since Perez’s automated volume, Ganguli’s human panels, and AISI’s threat-specific structured teams each found different things, and none of the three would have found what the others did. It would distinguish, in the same sentence, a production side-effect metric like a refusal-rate change from a red-team failure count like “no bypass in N hours” from a field-incident count, because these are different kinds of evidence and the rule-of-three bound only meaningfully applies to the second, and only once independent attempt counts, not hours, are disclosed. And it would treat a clean result as provisional rather than final, on a published re-audit schedule rather than at the lab’s discretion — the pattern CAISI and UK AISI’s standing testing agreements with five labs are already establishing, and the one the many-shot jailbreak and sleeper-agent findings show is necessary even absent any external adversary, since standard training can look like it worked without having worked at all.
For background, read the article’s source list.