Equation 11 · The Safeguards Behind a Claude Refusal, and What They're Actually Trading Off
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol τ
τ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Symbol P
P is an input to the expression that computes the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Every mechanism described above shares a single underlying structure, and naming it precisely is more useful than treating “safety” and “usefulness” as opposed virtues balanced by instinct. A classifier, or a trained refusal policy, reduces to a decision rule: given some score s(x) computed from a request — and, for output classifiers, a response — block if s(x) exceeds a threshold . Two error rates follow directly from that single rule: . Over-refusal is the false positive rate: benign requests wrongly blocked. Under-refusal is the false negative rate: harmful requests wrongly allowed through. Because both are read off the same score against the same threshold, they…
Read the full surrounding passage
Every mechanism described above shares a single underlying structure, and naming it precisely is more useful than treating “safety” and “usefulness” as opposed virtues balanced by instinct. A classifier, or a trained refusal policy, reduces to a decision rule: given some score s(x) computed from a request — and, for output classifiers, a response — block if s(x) exceeds a threshold . Two error rates follow directly from that single rule: . Over-refusal is the false positive rate: benign requests wrongly blocked. Under-refusal is the false negative rate: harmful requests wrongly allowed through. Because both are read off the same score against the same threshold, they move in opposite directions together — raising to let more borderline traffic through necessarily lets in more of the genuinely harmful traffic sitting just past that border, and lowering it to catch more harmful traffic necessarily catches more benign traffic that merely resembles it. This is a simplification of what Anthropic’s own cascade architecture actually does — a two-stage system with a cheap screening threshold and a separate, more careful second threshold is not fully captured by one scalar — but the simplification exposes the real constraint underneath the engineering: nothing described in this article moves one error rate without moving the other, and any published number for one without the other is only half a measurement.
Sources cited in the article section
- [10] XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models ↗
- [12] System Card: Claude Opus 4.5 ↗
- [5] Many-shot Jailbreaking ↗
These citations give research context. Read each source to check which claims it supports.
Return to The Safeguards Behind a Claude Refusal, and What They're Actually Trading Off