← All parts of this equation

Equation 1 · Part 7 · One Model, Many Modalities: What Multimodal Systems Actually Share

multiplication

Ntok=HP⋅WP,N_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P},
multiplication

What this part means

Multiply the quantities on either side.

Its job in the formula

Multiply the quantities on either side.

The passage around this formula

For images the dominant answer is patchification, introduced at scale by the Vision Transformer: cut the image into a grid of non-overlapping square patches, flatten each, and project it linearly into the model’s embedding dimension, so that a pure transformer applied directly to sequences of image patches performs competitively with convolutional networks when pre-trained on enough data [ 2 ] . The token count follows immediately from the geometry, Ntok=HP⋅WPN_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P}. for an image of height H and width W at patch size P . The quadratic relationship is the entire practical story of image tokenisation. Halving the patch size quadruples the sequence, and attention cost grows faster still.…

Read this part in the article →

Learn the underlying idea

Multiplication scales one quantity by another. A dot, a cross, or adjacent symbols can indicate a product.

Open the illustrated multiplication: combining factors guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.