← Back to article

Equation 1 · The Safeguards Behind a Claude Refusal, and What They're Actually Trading Off

What does this equation mean?

cprobec_{\text{probe}}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

the near-free cost of the first pass and cexchangec_{\text{exchange}} the much larger cost of the second. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

cprobec_{\text{probe}}

Symbol c_probe

the near-free cost of the first pass and cexchangec_{\text{exchange}} the much larger cost of the second.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

That first version cost real computation: 23.7% more inference compute than running the model alone, and a 0.38 percentage-point increase in refusals on ordinary, harmless traffic [ 4 ] . Anthropic’s January 2026 follow-up redesigned the system around a two-stage cascade rather than a single classifier pass: a lightweight probe reads the language model’s own internal activations and screens every exchange cheaply, escalating only the fraction it flags to a slower, more capable classifier that reviews the input and output together [ 1 ] . If cprobec_{\text{probe}} is the near-free cost of the first pass and cexchangec_{\text{exchange}} the much larger cost of the second, and pescp_{\text{esc}} is the fraction…
Read the full surrounding passage
That first version cost real computation: 23.7% more inference compute than running the model alone, and a 0.38 percentage-point increase in refusals on ordinary, harmless traffic [ 4 ] . Anthropic’s January 2026 follow-up redesigned the system around a two-stage cascade rather than a single classifier pass: a lightweight probe reads the language model’s own internal activations and screens every exchange cheaply, escalating only the fraction it flags to a slower, more capable classifier that reviews the input and output together [ 1 ] . If cprobec_{\text{probe}} is the near-free cost of the first pass and cexchangec_{\text{exchange}} the much larger cost of the second, and pescp_{\text{esc}} is the fraction of traffic the probe escalates, the expected added cost per request is

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to The Safeguards Behind a Claude Refusal, and What They're Actually Trading Off

Browse the mathematical compendium →