Equation 7 · What Multimodal AI Actually Costs, Modality by Modality
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol N_tok
ok is part of the quantity the equation computes from the expression on the right.
Symbol tau_frame
tarame is one of the signed contributions combined to compute the quantity on the left.
Symbol tau_audio
the per-frame and per-second-of-audio token costs Google publishes.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Video inherits the same per-tile arithmetic and multiplies it by time. Google is again the most specific vendor to check, because Gemini’s video-understanding documentation breaks the rate down by component: at default media resolution, each sampled frame (taken at one frame per second) costs about 258 tokens and the accompanying audio track costs about 32 tokens per second, for a documented total of “approximately 300 tokens per second of video.” At the model’s low-resolution setting, the same second costs about 100 tokens — 66 for the frame plus 32 for audio [ 6 ] . In the notation above, that is . where t is duration in seconds, f the sampling rate, and…
Read the full surrounding passage
Video inherits the same per-tile arithmetic and multiplies it by time. Google is again the most specific vendor to check, because Gemini’s video-understanding documentation breaks the rate down by component: at default media resolution, each sampled frame (taken at one frame per second) costs about 258 tokens and the accompanying audio track costs about 32 tokens per second, for a documented total of “approximately 300 tokens per second of video.” At the model’s low-resolution setting, the same second costs about 100 tokens — 66 for the frame plus 32 for audio [ 6 ] . In the notation above, that is . where t is duration in seconds, f the sampling rate, and , the per-frame and per-second-of-audio token costs Google publishes. A minute of default-resolution video comes to roughly 18,000 tokens — about 2.7 cents at Gemini 3.5 Flash’s input rate, or about 1.35 cents at Gemini 3.7 Flash’s [ 6 , 4 ] . That is understanding a minute of footage: having the model watch it once and answer questions about it.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to What Multimodal AI Actually Costs, Modality by Modality