← All parts of this equation

Equation 7 · Part 2 · What Multimodal AI Actually Costs, Modality by Modality

Symbol t

Ntok(video)=t(f⋅τframe+τaudio),N_{\mathrm{tok}}(\mathrm{video}) = t \left( f \cdot \tau_{\mathrm{frame}} + \tau_{\mathrm{audio}} \right),
tt

What this part means

duration in seconds.

Its job in the formula

t is one of the signed contributions combined to compute the quantity on the left.

Where the article explains it

where t is duration in seconds, f the sampling rate, and τframe\tau_{\mathrm{frame}} , τaudio\tau_{\mathrm{audio}} the per-frame and per-second-of-audio token costs Google publishes.

The passage around this formula

…300 tokens per second of video.” At the model’s low-resolution setting, the same second costs about 100 tokens — 66 for the frame plus 32 for audio [ 6 ] . In the notation above, that is Ntok(video)=t(f⋅τframe+τaudio)N_{\mathrm{tok}}(\mathrm{video}) = t \left( f \cdot \tau_{\mathrm{frame}} + \tau_{\mathrm{audio}} \right). where t is duration in seconds, f the sampling rate, and τframe\tau_{\mathrm{frame}} , τaudio\tau_{\mathrm{audio}} the per-frame and per-second-of-audio token costs Google publishes. A minute of default-resolution video comes to roughly 18,000 tokens — about 2.7 cents at Gemini 3.5 Flash’s input rate, or about 1.35…

Read this part in the article →

Learn the underlying idea

A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.

Open the illustrated functions: inputs become outputs guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.