← Back to article

Equation 12 · How Multimodal Models Actually Handle Video, Audio, and Space

What does this equation mean?

R=fsh⋅nq⋅log⁡2K bits per second.R = \frac{f_s}{h} \cdot n_q \cdot \log_2 K \text{ bits per second.}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Start withf_s
Divide byh
This relates toR
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

RR

Symbol R

the bitrate.

Understand this part →

fsf_s

Symbol f_s

fsf_s occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

hh

Symbol h

the number of samples.

Understand this part →

nqn_q

Symbol n_q

nqn_q is one factor in the product that computes the quantity on the left.

Understand this part →

KK

Symbol K

K is one factor in the product that computes the quantity on the left.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
multiplication

multiplication

Multiply the quantities on either side.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

The mechanism that does this at scale is a learned neural codec built around residual vector quantization: an encoder compresses the waveform into a sequence of continuous frames, and each frame is then quantized in successive stages, each stage encoding what the previous stage’s codebook missed. SoundStream is the reference architecture, described by its authors as an end-to-end neural audio codec built on a convolutional encoder-decoder pair with a residual vector quantizer, trained jointly and shown to outperform prior codecs at comparable bitrates [ 7 ] . The mechanism gives a concrete, computable token rate. For a signal sampled at fsf_s hertz, encoded with hop length h samples per frame,…
Read the full surrounding passage
The mechanism that does this at scale is a learned neural codec built around residual vector quantization: an encoder compresses the waveform into a sequence of continuous frames, and each frame is then quantized in successive stages, each stage encoding what the previous stage’s codebook missed. SoundStream is the reference architecture, described by its authors as an end-to-end neural audio codec built on a convolutional encoder-decoder pair with a residual vector quantizer, trained jointly and shown to outperform prior codecs at comparable bitrates [ 7 ] . The mechanism gives a concrete, computable token rate. For a signal sampled at fsf_s hertz, encoded with hop length h samples per frame, and quantized with nqn_q residual stages of codebook size K each, the bitrate is R=fsh⋅nq⋅log⁡2K bits per second.R = \frac{f_s}{h} \cdot n_q \cdot \log_2 K \text{ bits per second.}. Every added quantizer stage buys reconstruction fidelity at a fixed, computable token-rate cost — the audio-codec analogue of the video tokenizer’s frame-versus-resolution trade, and just as unavoidable.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to How Multimodal Models Actually Handle Video, Audio, and Space

See this formula across 1 published context →

Browse the mathematical compendium →