Foundation Models
The Safeguards Behind a Claude Refusal, and What They're Actually Trading Off
Every refusal passes through a trained policy, a classifier, account monitoring, and a capability-gated release process — four thresholds, each trading missed misuse against blocked legitimate use, documented in Anthropic's own safeguards and RSP filings.