The Hardest Unsolved Problems in Multimodal AI
Leaderboards read like steady progress toward one unified system. Five specific capabilities are not solved, by the researchers' own published admission — and each has a benchmark built to prove it.
Tagged multimodal AI · show all articles
Leaderboards read like steady progress toward one unified system. Five specific capabilities are not solved, by the researchers' own published admission — and each has a benchmark built to prove it.
Image understanding has become a converged, commodity layer across OpenAI, Google and Anthropic. Video and audio have not — and the documented architecture differences, not the marketing pages, show exactly where.
Two questions decide how multimodal AI matures by 2035: whether architecture converges on native any-to-any models, and whether perception and generation both cross into open-world reliability. Four scenarios, each with a falsifier.
Before a single model could read an image and write about it, two research communities spent a decade building separate pieces that had to be joined by hand. This is the record of each join.