← Mathematical compendium

Published equation contexts

Ntok=HP⋅WPN_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P}

Why this formula appears here

There is a structural reason this compounds rather than merely adding up, and it is worth making explicit because it exposes an assumption rather than a vague sense of difficulty. A vision transformer front end tokenises an image of height H and width W at patch size P into NtokN_{\mathrm{tok}} = HP\frac{H}{P} ⋅\cdot WP\frac{W}{P}, and a video of T sampled frames multiplies that by T again. Any fixed context budget therefore forces a three-way trade between spatial patch size, temporal sampling rate, and clip duration — coarsen the patches, subsample frames, or truncate the clip. Whichever is chosen, information is discarded before a language model ever sees a token, and the TemporalBench failure…

Read the full article-specific guide →

Read the representative guide

NtokN_{\mathrm{tok}}

Symbol N_tok

NtN_tok is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →
PP

Symbol P

P occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (3)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Ntok=HP⋅WP,N_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P},

Equation 4 · Foundation Models

The Hardest Unsolved Problems in Multimodal AI

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

There is a structural reason this compounds rather than merely adding up, and it is worth making explicit because it exposes an assumption rather than a vague sense of difficulty. A vision transformer front end tokenises an image of height H and width W at patch size P into NtokN_{\mathrm{tok}} = HP\frac{H}{P} ⋅\cdot WP\frac{W}{P}, and a video of T sampled frames multiplies that by T again. Any fixed context budget therefore forces a three-way trade between spatial patch size, temporal sampling rate, and clip duration — coarsen the patches, subsample frames, or truncate the clip. Whichever is chosen, information is discarded before a language model ever sees a token, and the TemporalBench failure…

Meanings in this article

  • HH: the height.
Equation guide → · Article →
Ntok=HP⋅WPN_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P}

Equation 4 · Foundation Models

How Multimodal Models Actually Handle Video, Audio, and Space

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

The image-text tokenization scheme has an exact, well-understood cost. An image of height H and width W , cut into square patches of size P , produces Ntok=HP⋅WPN_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P}. tokens. Video adds a third factor. If a model attends over T sampled frames, each cut into the same spatial grid, and groups PtP_t consecutive frames per temporal patch, the token count becomes

Meanings in this article

  • PP: the cut into square patches of size.
  • WW: the image of height h and width.
Equation guide → · Article →
Ntok=HP⋅WP,N_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P},

Equation 1 · Foundation Models

One Model, Many Modalities: What Multimodal Systems Actually Share

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

For images the dominant answer is patchification, introduced at scale by the Vision Transformer: cut the image into a grid of non-overlapping square patches, flatten each, and project it linearly into the model’s embedding dimension, so that a pure transformer applied directly to sequences of image patches performs competitively with convolutional networks when pre-trained on enough data [ 2 ] . The token count follows immediately from the geometry, Ntok=HP⋅WPN_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P}. for an image of height H and width W at patch size P . The quadratic relationship is the entire practical story of image tokenisation. Halving the patch size quadruples the sequence, and attention cost grows faster still.…

Meanings in this article

  • HH: the height.
Equation guide → · Article →