Equation 12 · The Safeguards Behind a Claude Refusal, and What They're Actually Trading Off
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol τ
τ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Over-refusal is the false positive rate: benign requests wrongly blocked. Under-refusal is the false negative rate: harmful requests wrongly allowed through. Because both are read off the same score against the same threshold, they move in opposite directions together — raising to let more borderline traffic through necessarily lets in more of the genuinely harmful traffic sitting just past that border, and lowering it to catch more harmful traffic necessarily catches more benign traffic that merely resembles it. This is a simplification of what Anthropic’s own cascade architecture actually does — a two-stage system with a cheap screening threshold and a separate, more careful second…
Read the full surrounding passage
Over-refusal is the false positive rate: benign requests wrongly blocked. Under-refusal is the false negative rate: harmful requests wrongly allowed through. Because both are read off the same score against the same threshold, they move in opposite directions together — raising to let more borderline traffic through necessarily lets in more of the genuinely harmful traffic sitting just past that border, and lowering it to catch more harmful traffic necessarily catches more benign traffic that merely resembles it. This is a simplification of what Anthropic’s own cascade architecture actually does — a two-stage system with a cheap screening threshold and a separate, more careful second threshold is not fully captured by one scalar — but the simplification exposes the real constraint underneath the engineering: nothing described in this article moves one error rate without moving the other, and any published number for one without the other is only half a measurement.
Sources cited in the article section
- [10] XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models ↗
- [12] System Card: Claude Opus 4.5 ↗
- [5] Many-shot Jailbreaking ↗
These citations give research context. Read each source to check which claims it supports.
Return to The Safeguards Behind a Claude Refusal, and What They're Actually Trading Off