Symbol N_tok
ok is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
There is a structural reason this compounds rather than merely adding up, and it is worth making explicit because it exposes an assumption rather than a vague sense of difficulty. A vision transformer front end tokenises an image of height H and width W at patch size P into = , and a video of T sampled frames multiplies that by T again. Any fixed context budget therefore forces a three-way trade between spatial patch size, temporal sampling rate, and clip duration — coarsen the patches, subsample frames, or truncate the clip. Whichever is chosen, information is discarded before a language model ever sees a token, and the TemporalBench failure…
ok is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →P occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Read this term in its guide →W occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Read this term in its guide →With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 4 · Foundation Models
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
There is a structural reason this compounds rather than merely adding up, and it is worth making explicit because it exposes an assumption rather than a vague sense of difficulty. A vision transformer front end tokenises an image of height H and width W at patch size P into = , and a video of T sampled frames multiplies that by T again. Any fixed context budget therefore forces a three-way trade between spatial patch size, temporal sampling rate, and clip duration — coarsen the patches, subsample frames, or truncate the clip. Whichever is chosen, information is discarded before a language model ever sees a token, and the TemporalBench failure…
Equation 4 · Foundation Models
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
The image-text tokenization scheme has an exact, well-understood cost. An image of height H and width W , cut into square patches of size P , produces . tokens. Video adds a third factor. If a model attends over T sampled frames, each cut into the same spatial grid, and groups consecutive frames per temporal patch, the token count becomes
Equation 1 · Foundation Models
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
For images the dominant answer is patchification, introduced at scale by the Vision Transformer: cut the image into a grid of non-overlapping square patches, flatten each, and project it linearly into the model’s embedding dimension, so that a pure transformer applied directly to sequences of image patches performs competitively with convolutional networks when pre-trained on enough data [ 2 ] . The token count follows immediately from the geometry, . for an image of height H and width W at patch size P . The quadratic relationship is the entire practical story of image tokenisation. Halving the patch size quadruples the sequence, and attention cost grows faster still.…