Equation 5 · How Do We Actually Know a Safety Measure Worked?
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol k
k is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Applied loosely to a red-team programme, this says a clean record is real evidence, and its strength depends entirely on k — the count of genuinely independent attempts, not the number of hours spent. The Constitutional Classifiers paper reports hours, not a count of independent attempts, so the bound cannot even be computed from the published disclosure as it stands, which is itself informative: hours are a proxy for effort, not a unit that determines a confidence level, and reporting the proxy rather than the underlying count is a gap in what the number can support. The approximation is also optimistic on its own terms, because real red-team attempts are not independent draws — they share…
Read the full surrounding passage
Applied loosely to a red-team programme, this says a clean record is real evidence, and its strength depends entirely on k — the count of genuinely independent attempts, not the number of hours spent. The Constitutional Classifiers paper reports hours, not a count of independent attempts, so the bound cannot even be computed from the published disclosure as it stands, which is itself informative: hours are a proxy for effort, not a unit that determines a confidence level, and reporting the proxy rather than the underlying count is a gap in what the number can support. The approximation is also optimistic on its own terms, because real red-team attempts are not independent draws — they share technique, tooling, and skill, so successive attempts by the same community are correlated rather than fresh trials, which pulls the true confidence bound weaker than the naive arithmetic suggests.
Sources cited in the article section
These citations give research context. Read each source to check which claims it supports.