Model Evaluation
Context Window Size Versus What a Frontier Model Can Actually Recall From It
OpenAI, Anthropic, and Google each publish a token count for how much a model can hold. Independent long-context benchmarks measure a different quantity: how much of it the model actually uses correctly, and the two numbers are routinely far apart.