← Mathematical compendium

Published equation contexts

Ntok(video)=t(f⋅τframe+τaudio)N_{\mathrm{tok}}(\mathrm{video}) = t \left( f \cdot \tau_{\mathrm{frame}} + \tau_{\mathrm{audio}} \right)

Why this formula appears here

Video inherits the same per-tile arithmetic and multiplies it by time. Google is again the most specific vendor to check, because Gemini’s video-understanding documentation breaks the rate down by component: at default media resolution, each sampled frame (taken at one frame per second) costs about 258 tokens and the accompanying audio track costs about 32 tokens per second, for a documented total of “approximately 300 tokens per second of video.” At the model’s low-resolution setting, the same second costs about 100 tokens — 66 for the frame plus 32 for audio [ 6 ] . In the notation above, that is Ntok(video)=t(f⋅τframe+τaudio)N_{\mathrm{tok}}(\mathrm{video}) = t \left( f \cdot \tau_{\mathrm{frame}} + \tau_{\mathrm{audio}} \right). where t is duration in seconds, f the sampling rate, and…

Read the full article-specific guide →

Read the representative guide

NtokN_{\mathrm{tok}}

Symbol N_tok

NtN_tok is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →
τframe\tau_{\mathrm{frame}}

Symbol tau_frame

taufu_frame is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Ntok(video)=t(f⋅τframe+τaudio),N_{\mathrm{tok}}(\mathrm{video}) = t \left( f \cdot \tau_{\mathrm{frame}} + \tau_{\mathrm{audio}} \right),

Equation 7 · Foundation Models

What Multimodal AI Actually Costs, Modality by Modality

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Video inherits the same per-tile arithmetic and multiplies it by time. Google is again the most specific vendor to check, because Gemini’s video-understanding documentation breaks the rate down by component: at default media resolution, each sampled frame (taken at one frame per second) costs about 258 tokens and the accompanying audio track costs about 32 tokens per second, for a documented total of “approximately 300 tokens per second of video.” At the model’s low-resolution setting, the same second costs about 100 tokens — 66 for the frame plus 32 for audio [ 6 ] . In the notation above, that is Ntok(video)=t(f⋅τframe+τaudio)N_{\mathrm{tok}}(\mathrm{video}) = t \left( f \cdot \tau_{\mathrm{frame}} + \tau_{\mathrm{audio}} \right). where t is duration in seconds, f the sampling rate, and…

Meanings in this article

Equation guide → · Article →