← Mathematical compendium

Published equation contexts

Ntok=HP⋅WP⋅TPtN_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P} \cdot \frac{T}{P_t}

Why this formula appears here

tokens. Video adds a third factor. If a model attends over T sampled frames, each cut into the same spatial grid, and groups PtP_t consecutive frames per temporal patch, the token count becomes Ntok=HP⋅WP⋅TPtN_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P} \cdot \frac{T}{P_t}. Every added frame multiplies the spatial cost rather than adding to it. A minute of video at even a modest frame rate is not “a bigger picture” in the way a higher-resolution photograph is; it is a volume, and attention cost over a transformer sequence grows faster than the sequence itself. No system deployed today attends densely over every raw frame of a long clip, because nothing could afford to.

Read the full article-specific guide →

Read the representative guide

NtokN_{\mathrm{tok}}

Symbol N_tok

NtN_tok is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →
PtP_t

Symbol P_t

PtP_t occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Ntok=HP⋅WP⋅TPt.N_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P} \cdot \frac{T}{P_t}.

Equation 7 · Foundation Models

How Multimodal Models Actually Handle Video, Audio, and Space

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

tokens. Video adds a third factor. If a model attends over T sampled frames, each cut into the same spatial grid, and groups PtP_t consecutive frames per temporal patch, the token count becomes Ntok=HP⋅WP⋅TPtN_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P} \cdot \frac{T}{P_t}. Every added frame multiplies the spatial cost rather than adding to it. A minute of video at even a modest frame rate is not “a bigger picture” in the way a higher-resolution photograph is; it is a volume, and attention cost over a transformer sequence grows faster than the sequence itself. No system deployed today attends densely over every raw frame of a long clip, because nothing could afford to.

Meanings in this article

  • PP: the cut into square patches of size.
  • WW: the image of height h and width.
Equation guide → · Article →