← All parts of this equation

Equation 7 · Part 9 · What Multimodal AI Actually Costs, Modality by Modality

subscript

Ntok(video)=t(f⋅τframe+τaudio),N_{\mathrm{tok}}(\mathrm{video}) = t \left( f \cdot \tau_{\mathrm{frame}} + \tau_{\mathrm{audio}} \right),
subscript

What this part means

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Its job in the formula

A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.

The passage around this formula

Video inherits the same per-tile arithmetic and multiplies it by time. Google is again the most specific vendor to check, because Gemini’s video-understanding documentation breaks the rate down by component: at default media resolution, each sampled frame (taken at one frame per second) costs about 258 tokens and the accompanying audio track costs about 32 tokens per second, for a documented total of “approximately 300 tokens per second of video.” At the model’s low-resolution setting, the same second costs about 100 tokens — 66 for the frame plus 32 for audio [ 6 ] . In the notation above, that is Ntok(video)=t(f⋅τframe+τaudio)N_{\mathrm{tok}}(\mathrm{video}) = t \left( f \cdot \tau_{\mathrm{frame}} + \tau_{\mathrm{audio}} \right). where t is duration in seconds, f the sampling rate, and…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.