Equation 4 · The Hardest Unsolved Problems in Multimodal AI
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol N_tok
ok is part of the quantity the equation computes from the expression on the right.
Symbol P
P occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Symbol W
W occurs above the fraction bar. The numerator is divided by the entire denominator below it.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
There is a structural reason this compounds rather than merely adding up, and it is worth making explicit because it exposes an assumption rather than a vague sense of difficulty. A vision transformer front end tokenises an image of height H and width W at patch size P into = , and a video of T sampled frames multiplies that by T again. Any fixed context budget therefore forces a three-way trade between spatial patch size, temporal sampling rate, and clip duration — coarsen the patches, subsample frames, or truncate the clip. Whichever is chosen, information is discarded before a language model ever sees a token, and the TemporalBench failure…
Read the full surrounding passage
There is a structural reason this compounds rather than merely adding up, and it is worth making explicit because it exposes an assumption rather than a vague sense of difficulty. A vision transformer front end tokenises an image of height H and width W at patch size P into = , and a video of T sampled frames multiplies that by T again. Any fixed context budget therefore forces a three-way trade between spatial patch size, temporal sampling rate, and clip duration — coarsen the patches, subsample frames, or truncate the clip. Whichever is chosen, information is discarded before a language model ever sees a token, and the TemporalBench failure mode — missing whether an action happened twice or three times — is exactly what temporal subsampling below the rate of the action would predict. This is not a claim that either paper makes explicitly; it is offered here as the mechanism that makes their independently reported findings consistent with each other rather than coincidental.
Sources cited in the article section
- [1] TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models ↗
- [2] LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding ↗
These citations give research context. Read each source to check which claims it supports.
Return to The Hardest Unsolved Problems in Multimodal AI