Equation 4 · Part 8 · How Multimodal Models Actually Handle Video, Audio, and Space
subscript
subscript
What this part means
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
Its job in the formula
A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.
Full expression→subscript→Article meaning
The passage around this formula
The image-text tokenization scheme has an exact, well-understood cost. An image of height H and width W , cut into square patches of size P , produces . tokens. Video adds a third factor. If a model attends over T sampled frames, each cut into the same spatial grid, and groups consecutive frames per temporal patch, the token count becomes
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
Sources cited in the article section
- [1] Video generation models as world simulators ↗
- [3] Language Model Beats Diffusion: Tokenizer is Key to Visual Generation ↗
- [2] VideoPoet: A Large Language Model for Zero-Shot Video Generation ↗
These citations provide research context; check each source for the exact claim it supports.