The Hardest Unsolved Problems in Multimodal AI
Leaderboards read like steady progress toward one unified system. Five specific capabilities are not solved, by the researchers' own published admission — and each has a benchmark built to prove it.
Tagged benchmark evaluation · show all articles
Leaderboards read like steady progress toward one unified system. Five specific capabilities are not solved, by the researchers' own published admission — and each has a benchmark built to prove it.
A score like SWE-bench 74 percent looks like a fact about a model. It is really a fact about a model, a harness, a task subset, and a date bundled together, and mixing those up is the most common way to misread a benchmark.
A model that sees, hears and reads still returns one fluent answer, and fluency hides the fact that the link between that answer and what was actually there breaks in ten distinct, independently documented ways.
Two variables decide the frontier-model contest by 2035: how many labs stay near the frontier, and whether the public signal of leadership is benchmark score or deployment outcome. Four scenarios follow, each with a falsifier.