← Back to article

Equation 4 · The Safeguards Behind a Claude Refusal, and What They're Actually Trading Off

What does this equation mean?

cˉ=cprobe+pesc⋅cexchange.\bar{c} = c_{\text{probe}} + p_{\text{esc}} \cdot c_{\text{exchange}}.

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationsc_probe + p_esc × c_exchange
Result or conditionbarc
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

cˉ\bar{c}

Symbol barc

the expected added cost per request.

Understand this part →

cprobec_{\text{probe}}

Symbol c_probe

the near-free cost of the first pass and cexchangec_{\text{exchange}} the much larger cost of the second.

Understand this part →

pescp_{\text{esc}}

Symbol p_esc

pep_esc is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

cexchangec_{\text{exchange}}

Symbol c_exchange

the much larger cost of the second.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
multiplication

multiplication

Multiply the quantities on either side.

Understand this part →

addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

That first version cost real computation: 23.7% more inference compute than running the model alone, and a 0.38 percentage-point increase in refusals on ordinary, harmless traffic [ 4 ] . Anthropic’s January 2026 follow-up redesigned the system around a two-stage cascade rather than a single classifier pass: a lightweight probe reads the language model’s own internal activations and screens every exchange cheaply, escalating only the fraction it flags to a slower, more capable classifier that reviews the input and output together [ 1 ] . If cprobec_{\text{probe}} is the near-free cost of the first pass and cexchangec_{\text{exchange}} the much larger cost of the second, and pescp_{\text{esc}} is the fraction…
Read the full surrounding passage
That first version cost real computation: 23.7% more inference compute than running the model alone, and a 0.38 percentage-point increase in refusals on ordinary, harmless traffic [ 4 ] . Anthropic’s January 2026 follow-up redesigned the system around a two-stage cascade rather than a single classifier pass: a lightweight probe reads the language model’s own internal activations and screens every exchange cheaply, escalating only the fraction it flags to a slower, more capable classifier that reviews the input and output together [ 1 ] . If cprobec_{\text{probe}} is the near-free cost of the first pass and cexchangec_{\text{exchange}} the much larger cost of the second, and pescp_{\text{esc}} is the fraction of traffic the probe escalates, the expected added cost per request is cˉ=cprobe+pesc⋅cexchange\bar{c} = c_{\text{probe}} + p_{\text{esc}} \cdot c_{\text{exchange}}. Because pescp_{\text{esc}} is small on ordinary traffic, the average cost collapses toward cprobec_{\text{probe}} even though cexchangec_{\text{exchange}} did not get any cheaper — which is the documented reason Anthropic reports the overhead falling from 23.7% to roughly 1% between generations, alongside a drop in the harmless-traffic refusal rate to 0.05% and, across more than 1,700 cumulative red-teaming hours against the new system, one high-risk vulnerability found, which Anthropic describes as a detection rate of 0.005 per thousand queries — the lowest of any defense it reports having evaluated [ 1 ] . Every one of those figures is Anthropic’s own measurement of its own system; none has yet appeared in an independently authored, peer-reviewed replication.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to The Safeguards Behind a Claude Refusal, and What They're Actually Trading Off

See this formula across 1 published context →

Browse the mathematical compendium →