← All parts of this equation

Equation 11 · Part 1 · The Safeguards Behind a Claude Refusal, and What They're Actually Trading Off

Symbol τ

FPR(τ)=P(block∣benign),FNR(τ)=P(allow∣harmful).\text{FPR}(\tau) = P(\text{block} \mid \text{benign}), \qquad \text{FNR}(\tau) = P(\text{allow} \mid \text{harmful}).
τ\tau

What this part means

τ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Its job in the formula

τ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

The passage around this formula

…balanced by instinct. A classifier, or a trained refusal policy, reduces to a decision rule: given some score s(x) computed from a request — and, for output classifiers, a response — block if s(x) exceeds a threshold τ\tau . Two error rates follow directly from that single rule: FPR(τ)=P(block∣benign),FNR(τ)=P(allow∣harmful)\text{FPR}(\tau) = P(\text{block} \mid \text{benign}), \qquad \text{FNR}(\tau) = P(\text{allow} \mid \text{harmful}). Over-refusal is the false positive rate: benign requests wrongly blocked. Under-refusal is the false negative rate: harmful requests wrongly allowed through. Because both are read off the same score against the same…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.