← All parts of this equation

Equation 4 · Part 2 · Where Multimodal Frontier Models Actually Differ, Beyond the Marketing

≈

τaudio≈32\tau_{\text{audio}} \approx 32
≈

What this part means

Approximately equal to; the equality is not exact.

Its job in the formula

Approximately equal to; the equality is not exact.

The passage around this formula

where, at default resolution, the sampling rate f is one frame per second, each frame costs τframe\tau_{\text{frame}} ≈\approx 258 tokens, and the audio track costs τaudio\tau_{\text{audio}} ≈\approx 32 tokens per second — giving cvideoc_{\text{video}} ≈\approx 290 tokens per second, close to the approximately 300 tokens per second Google documents directly for default-resolution video [ 5 ] . The formula is unremarkable arithmetic, but what it exposes is the real assumption underneath the “native” claim: a single request is billed, and therefore presumably processed, as one combined audio-visual stream, not as a picture track with a separately bolted-on transcript.

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.