← All parts of this equation

Equation 7 · Part 6 · How Multimodal Models Actually Handle Video, Audio, and Space

Symbol P_t

Ntok=HP⋅WP⋅TPt.N_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P} \cdot \frac{T}{P_t}.
PtP_t

What this part means

PtP_t occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Its job in the formula

PtP_t occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

The passage around this formula

tokens. Video adds a third factor. If a model attends over T sampled frames, each cut into the same spatial grid, and groups PtP_t consecutive frames per temporal patch, the token count becomes Ntok=HP⋅WP⋅TPtN_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P} \cdot \frac{T}{P_t}. Every added frame multiplies the spatial cost rather than adding to it. A minute of video at even a modest frame rate is not “a bigger picture” in the way a higher-resolution photograph is; it is a volume, and attention cost over a transformer sequence grows faster than the sequence itself. No system deployed today attends densely over every raw frame of a long clip, because nothing could afford to.

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.