Model Evaluation
OpenAI and Claude on Formal Reasoning: What the Benchmarks Show, and Where They Mislead
AIME, GPQA and FrontierMath test different things by different rules. Reasoning effort and extended thinking are documented settings, not fixed traits, and a right answer does not certify the reasoning behind it.