← All parts of this equation

Equation 4 · Part 1 · Where Multimodal Frontier Models Actually Differ, Beyond the Marketing

Symbol tau_audio

τaudio≈32\tau_{\text{audio}} \approx 32
τaudio\tau_{\text{audio}}

What this part means

tauau_audio is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Its job in the formula

tauau_audio is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

The passage around this formula

where, at default resolution, the sampling rate f is one frame per second, each frame costs τframe\tau_{\text{frame}} ≈\approx 258 tokens, and the audio track costs τaudio\tau_{\text{audio}} ≈\approx 32 tokens per second — giving cvideoc_{\text{video}} ≈\approx 290 tokens per second, close to the approximately 300 tokens per second Google documents directly for default-resolution video [ 5 ] . The formula is unremarkable arithmetic, but what it exposes is the real assumption underneath the “native” claim: a single request is billed, and therefore presumably processed, as one combined audio-visual stream, not as a picture track with a separately bolted-on transcript.

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.