← All parts of this equation

Equation 12 · Part 1 · How Multimodal Models Actually Handle Video, Audio, and Space

Symbol R

R=fsh⋅nq⋅log⁡2K bits per second.R = \frac{f_s}{h} \cdot n_q \cdot \log_2 K \text{ bits per second.}
RR

What this part means

the bitrate.

Its job in the formula

R is part of the quantity the equation computes from the expression on the right.

Where the article explains it

For a signal sampled at fsf_s hertz, encoded with hop length h samples per frame, and quantized with nqn_q residual stages of codebook size K each, the bitrate is

The passage around this formula

The mechanism that does this at scale is a learned neural codec built around residual vector quantization: an encoder compresses the waveform into a sequence of continuous frames, and each frame is then quantized in successive stages, each stage encoding what the previous stage’s codebook missed. SoundStream is the reference architecture, described by its authors as an end-to-end neural audio codec built on a convolutional encoder-decoder pair with a residual vector quantizer, trained jointly and shown to outperform prior codecs at comparable bitrates [ 7 ] . The mechanism gives a concrete, computable token rate. For a signal sampled at fsf_s hertz, encoded with hop length h samples per frame,…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.