← All parts of this equation

Equation 11 · Part 2 · The Safeguards Behind a Claude Refusal, and What They're Actually Trading Off

Symbol P

FPR(τ)=P(block∣benign),FNR(τ)=P(allow∣harmful).\text{FPR}(\tau) = P(\text{block} \mid \text{benign}), \qquad \text{FNR}(\tau) = P(\text{allow} \mid \text{harmful}).
PP

What this part means

P is an input to the expression that computes the quantity on the left.

Its job in the formula

P is an input to the expression that computes the quantity on the left.

The passage around this formula

Every mechanism described above shares a single underlying structure, and naming it precisely is more useful than treating “safety” and “usefulness” as opposed virtues balanced by instinct. A classifier, or a trained refusal policy, reduces to a decision rule: given some score s(x) computed from a request — and, for output classifiers, a response — block if s(x) exceeds a threshold τ\tau . Two error rates follow directly from that single rule: FPR(τ)=P(block∣benign),FNR(τ)=P(allow∣harmful)\text{FPR}(\tau) = P(\text{block} \mid \text{benign}), \qquad \text{FNR}(\tau) = P(\text{allow} \mid \text{harmful}). Over-refusal is the false positive rate: benign requests wrongly blocked. Under-refusal is the false negative rate: harmful requests wrongly allowed through. Because both are read off the same score against the same threshold, they…

Read this part in the article →

Learn the underlying idea

A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.

Open the illustrated functions: inputs become outputs guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.