← All parts of this equation

Equation 7 · Part 1 · What Multimodal AI Actually Costs, Modality by Modality

Symbol N_tok

Ntok(video)=t(f⋅τframe+τaudio),N_{\mathrm{tok}}(\mathrm{video}) = t \left( f \cdot \tau_{\mathrm{frame}} + \tau_{\mathrm{audio}} \right),
NtokN_{\mathrm{tok}}

What this part means

NtN_tok is part of the quantity the equation computes from the expression on the right.

Its job in the formula

NtN_tok is part of the quantity the equation computes from the expression on the right.

The passage around this formula

Video inherits the same per-tile arithmetic and multiplies it by time. Google is again the most specific vendor to check, because Gemini’s video-understanding documentation breaks the rate down by component: at default media resolution, each sampled frame (taken at one frame per second) costs about 258 tokens and the accompanying audio track costs about 32 tokens per second, for a documented total of “approximately 300 tokens per second of video.” At the model’s low-resolution setting, the same second costs about 100 tokens — 66 for the frame plus 32 for audio [ 6 ] . In the notation above, that is Ntok(video)=t(f⋅τframe+τaudio)N_{\mathrm{tok}}(\mathrm{video}) = t \left( f \cdot \tau_{\mathrm{frame}} + \tau_{\mathrm{audio}} \right). where t is duration in seconds, f the sampling rate, and…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.