Symbol N_tok
ok is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
tokens. Video adds a third factor. If a model attends over T sampled frames, each cut into the same spatial grid, and groups consecutive frames per temporal patch, the token count becomes . Every added frame multiplies the spatial cost rather than adding to it. A minute of video at even a modest frame rate is not “a bigger picture” in the way a higher-resolution photograph is; it is a volume, and attention cost over a transformer sequence grows faster than the sequence itself. No system deployed today attends densely over every raw frame of a long clip, because nothing could afford to.
ok is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →H occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Read this term in its guide →T occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Read this term in its guide →occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Read this term in its guide →With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 7 · Foundation Models
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
tokens. Video adds a third factor. If a model attends over T sampled frames, each cut into the same spatial grid, and groups consecutive frames per temporal patch, the token count becomes . Every added frame multiplies the spatial cost rather than adding to it. A minute of video at even a modest frame rate is not “a bigger picture” in the way a higher-resolution photograph is; it is a volume, and attention cost over a transformer sequence grows faster than the sequence itself. No system deployed today attends densely over every raw frame of a long clip, because nothing could afford to.