Equation 4 · The Safeguards Behind a Claude Refusal, and What They're Actually Trading Off
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol c_probe
the near-free cost of the first pass and the much larger cost of the second.
Symbol p_esc
sc is one of the signed contributions combined to compute the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
That first version cost real computation: 23.7% more inference compute than running the model alone, and a 0.38 percentage-point increase in refusals on ordinary, harmless traffic [ 4 ] . Anthropic’s January 2026 follow-up redesigned the system around a two-stage cascade rather than a single classifier pass: a lightweight probe reads the language model’s own internal activations and screens every exchange cheaply, escalating only the fraction it flags to a slower, more capable classifier that reviews the input and output together [ 1 ] . If is the near-free cost of the first pass and the much larger cost of the second, and is the fraction…
Read the full surrounding passage
That first version cost real computation: 23.7% more inference compute than running the model alone, and a 0.38 percentage-point increase in refusals on ordinary, harmless traffic [ 4 ] . Anthropic’s January 2026 follow-up redesigned the system around a two-stage cascade rather than a single classifier pass: a lightweight probe reads the language model’s own internal activations and screens every exchange cheaply, escalating only the fraction it flags to a slower, more capable classifier that reviews the input and output together [ 1 ] . If is the near-free cost of the first pass and the much larger cost of the second, and is the fraction of traffic the probe escalates, the expected added cost per request is . Because is small on ordinary traffic, the average cost collapses toward even though did not get any cheaper — which is the documented reason Anthropic reports the overhead falling from 23.7% to roughly 1% between generations, alongside a drop in the harmless-traffic refusal rate to 0.05% and, across more than 1,700 cumulative red-teaming hours against the new system, one high-risk vulnerability found, which Anthropic describes as a detection rate of 0.005 per thousand queries — the lowest of any defense it reports having evaluated [ 1 ] . Every one of those figures is Anthropic’s own measurement of its own system; none has yet appeared in an independently authored, peer-reviewed replication.
Sources cited in the surrounding passage
- [4] Constitutional Classifiers: Defending Against Universal Jailbreaks across Thousands of Hours of Red Teaming ↗
- [1] Next-Generation Constitutional Classifiers: More Efficient Protection Against Universal Jailbreaks ↗
These citations give research context. Read each source to check which claims it supports.
Return to The Safeguards Behind a Claude Refusal, and What They're Actually Trading Off