Published equation contexts
Why this formula appears here
The policy’s structure is what makes it useful to state formally, because its core mechanism is a conditional commitment rather than a fixed rule. Anthropic defines AI Safety Levels (ASL), each tied to a capability threshold on a specific class of risk — most concretely, the potential for a model to meaningfully assist in acquiring chemical, biological, radiological, or nuclear weapons capability, or to self-exfiltrate or resist correction. The policy’s operating logic, stripped to its structural claim, is a threshold trigger: . where C() is the model’s evaluated capability on a specific threat class, is the threshold defining ASL level k , and is…
Read the representative guide
Symbol S_k
the security and deployment standard the policy requires once that threshold is crossed.
Read this term in its guide →How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
Research cited beside this formula
Published contexts (1)
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 1 · Foundation Models
A History of Anthropic and the Claude Model Line
This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions.
The policy’s structure is what makes it useful to state formally, because its core mechanism is a conditional commitment rather than a fixed rule. Anthropic defines AI Safety Levels (ASL), each tied to a capability threshold on a specific class of risk — most concretely, the potential for a model to meaningfully assist in acquiring chemical, biological, radiological, or nuclear weapons capability, or to self-exfiltrate or resist correction. The policy’s operating logic, stripped to its structural claim, is a threshold trigger: . where C() is the model’s evaluated capability on a specific threat class, is the threshold defining ASL level k , and is…
Meanings in this article
- : the model’s evaluated capability on a specific threat class.
- : the threshold defining ASL level k.
- : the security and deployment standard the policy requires once that threshold is crossed.