Equation 4 · Part 2 · The Hardest Unsolved Problems in Multimodal AI
Symbol H
What this part means
the height.
Its job in the formula
H occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Full expression→Symbol H→Article meaning
Where the article explains it
A vision transformer front end tokenises an image of height H and width W at patch size P into = , and a video of T sampled frames multiplies that by T again.
The passage around this formula
…reason this compounds rather than merely adding up, and it is worth making explicit because it exposes an assumption rather than a vague sense of difficulty. A vision transformer front end tokenises an image of height H and width W at patch size P into = , and a video of T sampled frames multiplies that by T again. Any fixed context budget therefore forces a three-way trade between spatial patch size, temporal sampling rate, and clip duration — coarsen the patches,…
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the article section
- [1] TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models ↗
- [2] LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding ↗
These citations provide research context; check each source for the exact claim it supports.