← All parts of this equation

Equation 7 · Part 8 · What Multimodal AI Actually Costs, Modality by Modality

addition

Ntok(video)=t(f⋅τframe+τaudio),N_{\mathrm{tok}}(\mathrm{video}) = t \left( f \cdot \tau_{\mathrm{frame}} + \tau_{\mathrm{audio}} \right),
addition

What this part means

Add the term after the plus sign to the term or group before it.

Its job in the formula

Add the term after the plus sign to the term or group before it.

The passage around this formula

Video inherits the same per-tile arithmetic and multiplies it by time. Google is again the most specific vendor to check, because Gemini’s video-understanding documentation breaks the rate down by component: at default media resolution, each sampled frame (taken at one frame per second) costs about 258 tokens and the accompanying audio track costs about 32 tokens per second, for a documented total of “approximately 300 tokens per second of video.” At the model’s low-resolution setting, the same second costs about 100 tokens — 66 for the frame plus 32 for audio [ 6 ] . In the notation above, that is Ntok(video)=t(f⋅τframe+τaudio)N_{\mathrm{tok}}(\mathrm{video}) = t \left( f \cdot \tau_{\mathrm{frame}} + \tau_{\mathrm{audio}} \right). where t is duration in seconds, f the sampling rate, and…

Read this part in the article →

Learn the underlying idea

Addition combines quantities; subtraction measures the signed difference between them. Parentheses show what is combined before the rest of the expression is evaluated.

Open the illustrated addition and subtraction in an equation guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.