Equation 7 · Part 3 · How Multimodal Models Actually Handle Video, Audio, and Space
Symbol P
What this part means
the cut into square patches of size.
Its job in the formula
P occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Full expression→Symbol P→Article meaning
Where the article explains it
An image of height H and width W , cut into square patches of size P , produces = tokens.
The passage around this formula
tokens. Video adds a third factor. If a model attends over T sampled frames, each cut into the same spatial grid, and groups consecutive frames per temporal patch, the token count becomes . Every added frame multiplies the spatial cost rather than adding to it. A minute of video at even a modest frame rate is not “a bigger picture” in the way a higher-resolution photograph is; it is a volume, and attention cost over a transformer sequence grows faster than the sequence itself. No system deployed today attends densely over every raw frame of a long clip, because nothing could afford to.
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the article section
- [1] Video generation models as world simulators ↗
- [3] Language Model Beats Diffusion: Tokenizer is Key to Visual Generation ↗
- [2] VideoPoet: A Large Language Model for Zero-Shot Video Generation ↗
These citations provide research context; check each source for the exact claim it supports.