Equation 7 · How Multimodal Models Actually Handle Video, Audio, and Space
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol N_tok
ok is part of the quantity the equation computes from the expression on the right.
Symbol H
H occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Symbol T
T occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Symbol P_t
occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
tokens. Video adds a third factor. If a model attends over T sampled frames, each cut into the same spatial grid, and groups consecutive frames per temporal patch, the token count becomes . Every added frame multiplies the spatial cost rather than adding to it. A minute of video at even a modest frame rate is not “a bigger picture” in the way a higher-resolution photograph is; it is a volume, and attention cost over a transformer sequence grows faster than the sequence itself. No system deployed today attends densely over every raw frame of a long clip, because nothing could afford to.
Sources cited in the article section
- [1] Video generation models as world simulators ↗
- [3] Language Model Beats Diffusion: Tokenizer is Key to Visual Generation ↗
- [2] VideoPoet: A Large Language Model for Zero-Shot Video Generation ↗
These citations give research context. Read each source to check which claims it supports.
Return to How Multimodal Models Actually Handle Video, Audio, and Space