Equation 7 · Part 9 · How Multimodal Models Actually Handle Video, Audio, and Space
multiplication
multiplication
What this part means
Multiply the quantities on either side.
Its job in the formula
Multiply the quantities on either side.
Full expression→multiplication→Article meaning
The passage around this formula
tokens. Video adds a third factor. If a model attends over T sampled frames, each cut into the same spatial grid, and groups consecutive frames per temporal patch, the token count becomes . Every added frame multiplies the spatial cost rather than adding to it. A minute of video at even a modest frame rate is not “a bigger picture” in the way a higher-resolution photograph is; it is a volume, and attention cost over a transformer sequence grows faster than the sequence itself. No system deployed today attends densely over every raw frame of a long clip, because nothing could afford to.
Learn the underlying idea
Multiplication scales one quantity by another. A dot, a cross, or adjacent symbols can indicate a product.
Open the illustrated multiplication: combining factors guide →
Sources cited in the article section
- [1] Video generation models as world simulators ↗
- [3] Language Model Beats Diffusion: Tokenizer is Key to Visual Generation ↗
- [2] VideoPoet: A Large Language Model for Zero-Shot Video Generation ↗
These citations provide research context; check each source for the exact claim it supports.