← All parts of this equation

Equation 4 · Part 3 · How Multimodal Models Actually Handle Video, Audio, and Space

Symbol P

Ntok=HP⋅WPN_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P}
PP

What this part means

the cut into square patches of size.

Its job in the formula

P occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Where the article explains it

An image of height H and width W , cut into square patches of size P , produces Ntok=HP⋅WPN_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P}.

The passage around this formula

The image-text tokenization scheme has an exact, well-understood cost. An image of height H and width W , cut into square patches of size P , produces Ntok=HP⋅WPN_{\mathrm{tok}} = \frac{H}{P} \cdot \frac{W}{P}. tokens. Video adds a third factor. If a model attends over T sampled frames, each cut into the same spatial grid, and groups PtP_t consecutive frames per temporal patch, the token count becomes

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.