The Hardest Unsolved Problems in Multimodal AI
Leaderboards read like steady progress toward one unified system. Five specific capabilities are not solved, by the researchers' own published admission — and each has a benchmark built to prove it.
Tagged video understanding · show all articles
Leaderboards read like steady progress toward one unified system. Five specific capabilities are not solved, by the researchers' own published admission — and each has a benchmark built to prove it.
Image understanding has become a converged, commodity layer across OpenAI, Google and Anthropic. Video and audio have not — and the documented architecture differences, not the marketing pages, show exactly where.
Image-text systems tokenize a grid and align two vectors. Video is a spacetime volume, audio is a waveform, and a 3D scene is a continuous field — none of them arrives as a grid, and each forces its own discretization scheme before any model can read it.