← All parts of this equation

Equation 7 · Part 5 · How Multimodal Models Actually Handle Video, Audio, and Space

Symbol T

Ntok=HP⋅WP⋅TPt.N_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P} \cdot \frac{T}{P_t}.
TT

What this part means

T occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Its job in the formula

T occurs above the fraction bar. The numerator is divided by the entire denominator below it.

The passage around this formula

tokens. Video adds a third factor. If a model attends over T sampled frames, each cut into the same spatial grid, and groups PtP_t consecutive frames per temporal patch, the token count becomes Ntok=HP⋅WP⋅TPtN_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P} \cdot \frac{T}{P_t}. Every added frame multiplies the spatial cost rather than adding to it. A minute of video at even a modest frame rate is not “a bigger picture” in the way a higher-resolution photograph is; it is a volume, and attention cost over a transformer sequence grows faster than the sequence itself. No system deployed today attends densely over every raw frame of a long clip, because nothing could afford to.

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.