← All parts of this equation

Equation 7 · Part 8 · How Multimodal Models Actually Handle Video, Audio, and Space

fraction

Ntok=HP⋅WP⋅TPt.N_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P} \cdot \frac{T}{P_t}.
fraction

What this part means

Divide the expression above the line by the one below it.

Its job in the formula

The expression above the fraction bar is divided by the complete expression below it. The denominator must not be zero.

The passage around this formula

tokens. Video adds a third factor. If a model attends over T sampled frames, each cut into the same spatial grid, and groups PtP_t consecutive frames per temporal patch, the token count becomes Ntok=HP⋅WP⋅TPtN_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P} \cdot \frac{T}{P_t}. Every added frame multiplies the spatial cost rather than adding to it. A minute of video at even a modest frame rate is not “a bigger picture” in the way a higher-resolution photograph is; it is a volume, and attention cost over a transformer sequence grows faster than the sequence itself. No system deployed today attends densely over every raw frame of a long clip, because nothing could afford to.

Read this part in the article →

Learn the underlying idea

A fraction a/b means a divided by b. The top number is the numerator; the bottom number is the denominator, and it cannot be zero.

Open the illustrated fractions: division written vertically guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.