← Mathematical compendium

Published equation contexts

FPR(τ)=P(block∣benign),FNR(τ)=P(allow∣harmful)\text{FPR}(\tau) = P(\text{block} \mid \text{benign}), \qquad \text{FNR}(\tau) = P(\text{allow} \mid \text{harmful})

Why this formula appears here

Every mechanism described above shares a single underlying structure, and naming it precisely is more useful than treating “safety” and “usefulness” as opposed virtues balanced by instinct. A classifier, or a trained refusal policy, reduces to a decision rule: given some score s(x) computed from a request — and, for output classifiers, a response — block if s(x) exceeds a threshold τ\tau . Two error rates follow directly from that single rule: FPR(τ)=P(block∣benign),FNR(τ)=P(allow∣harmful)\text{FPR}(\tau) = P(\text{block} \mid \text{benign}), \qquad \text{FNR}(\tau) = P(\text{allow} \mid \text{harmful}). Over-refusal is the false positive rate: benign requests wrongly blocked. Under-refusal is the false negative rate: harmful requests wrongly allowed through. Because both are read off the same score against the same threshold, they…

Read the full article-specific guide →

Read the representative guide

τ\tau

Symbol τ

τ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

FPR(τ)=P(block∣benign),FNR(τ)=P(allow∣harmful).\text{FPR}(\tau) = P(\text{block} \mid \text{benign}), \qquad \text{FNR}(\tau) = P(\text{allow} \mid \text{harmful}).

Equation 11 · Foundation Models

The Safeguards Behind a Claude Refusal, and What They're Actually Trading Off

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Every mechanism described above shares a single underlying structure, and naming it precisely is more useful than treating “safety” and “usefulness” as opposed virtues balanced by instinct. A classifier, or a trained refusal policy, reduces to a decision rule: given some score s(x) computed from a request — and, for output classifiers, a response — block if s(x) exceeds a threshold τ\tau . Two error rates follow directly from that single rule: FPR(τ)=P(block∣benign),FNR(τ)=P(allow∣harmful)\text{FPR}(\tau) = P(\text{block} \mid \text{benign}), \qquad \text{FNR}(\tau) = P(\text{allow} \mid \text{harmful}). Over-refusal is the false positive rate: benign requests wrongly blocked. Under-refusal is the false negative rate: harmful requests wrongly allowed through. Because both are read off the same score against the same threshold, they…

Equation guide → · Article →