← Back to article

Equation 11 · The Safeguards Behind a Claude Refusal, and What They're Actually Trading Off

What does this equation mean?

FPR(τ)=P(block∣benign),FNR(τ)=P(allow∣harmful).\text{FPR}(\tau) = P(\text{block} \mid \text{benign}), \qquad \text{FNR}(\tau) = P(\text{allow} \mid \text{harmful}).

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationsP(block mid benign), qquad FNR(τ) = P(allow mid harmful)
Result or conditionFPR(τ)
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

τ\tau

Symbol τ

τ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

PP

Symbol P

P is an input to the expression that computes the quantity on the left.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Every mechanism described above shares a single underlying structure, and naming it precisely is more useful than treating “safety” and “usefulness” as opposed virtues balanced by instinct. A classifier, or a trained refusal policy, reduces to a decision rule: given some score s(x) computed from a request — and, for output classifiers, a response — block if s(x) exceeds a threshold τ\tau . Two error rates follow directly from that single rule: FPR(τ)=P(block∣benign),FNR(τ)=P(allow∣harmful)\text{FPR}(\tau) = P(\text{block} \mid \text{benign}), \qquad \text{FNR}(\tau) = P(\text{allow} \mid \text{harmful}). Over-refusal is the false positive rate: benign requests wrongly blocked. Under-refusal is the false negative rate: harmful requests wrongly allowed through. Because both are read off the same score against the same threshold, they…
Read the full surrounding passage
Every mechanism described above shares a single underlying structure, and naming it precisely is more useful than treating “safety” and “usefulness” as opposed virtues balanced by instinct. A classifier, or a trained refusal policy, reduces to a decision rule: given some score s(x) computed from a request — and, for output classifiers, a response — block if s(x) exceeds a threshold τ\tau . Two error rates follow directly from that single rule: FPR(τ)=P(block∣benign),FNR(τ)=P(allow∣harmful)\text{FPR}(\tau) = P(\text{block} \mid \text{benign}), \qquad \text{FNR}(\tau) = P(\text{allow} \mid \text{harmful}). Over-refusal is the false positive rate: benign requests wrongly blocked. Under-refusal is the false negative rate: harmful requests wrongly allowed through. Because both are read off the same score against the same threshold, they move in opposite directions together — raising τ\tau to let more borderline traffic through necessarily lets in more of the genuinely harmful traffic sitting just past that border, and lowering it to catch more harmful traffic necessarily catches more benign traffic that merely resembles it. This is a simplification of what Anthropic’s own cascade architecture actually does — a two-stage system with a cheap screening threshold and a separate, more careful second threshold is not fully captured by one scalar τ\tau — but the simplification exposes the real constraint underneath the engineering: nothing described in this article moves one error rate without moving the other, and any published number for one without the other is only half a measurement.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to The Safeguards Behind a Claude Refusal, and What They're Actually Trading Off

See this formula across 1 published context →

Browse the mathematical compendium →