Model Evaluation
Two Different Bets on How to Align a Frontier Model
OpenAI and Anthropic both call their models "aligned." The mechanisms behind that word, the published thresholds for when to stop scaling, and what independent red-teamers find when they push on both are three separate questions with three separate answers.