Foundation Models
What Claude's Capability Evaluations Actually Measure
A benchmark score is not a fact about a model; it is a fact about a harness, a threshold, and a moment in time. Here is what actually goes into an Anthropic capability evaluation, who checks it from outside, and why a single leaderboard number still misleads.