← All parts of this equation

Equation 12 · Part 3 · How Multimodal Models Actually Handle Video, Audio, and Space

Symbol h

R=fsh⋅nq⋅log⁡2K bits per second.R = \frac{f_s}{h} \cdot n_q \cdot \log_2 K \text{ bits per second.}
hh

What this part means

the number of samples.

Its job in the formula

h occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Where the article explains it

For a signal sampled at fsf_s hertz, encoded with hop length h samples per frame, and quantized with nqn_q residual stages of codebook size K each, the bitrate is R=fsh⋅nq⋅log⁡2K bits per second.R = \frac{f_s}{h} \cdot n_q \cdot \log_2 K \text{ bits per second.}.

The passage around this formula

…vector quantizer, trained jointly and shown to outperform prior codecs at comparable bitrates [ 7 ] . The mechanism gives a concrete, computable token rate. For a signal sampled at fsf_s hertz, encoded with hop length h samples per frame, and quantized with nqn_q residual stages of codebook size K each, the bitrate is R=fsh⋅nq⋅log⁡2K bits per second.R = \frac{f_s}{h} \cdot n_q \cdot \log_2 K \text{ bits per second.}. Every added quantizer stage buys reconstruction fidelity at a fixed, computable token-rate cost — the audio-codec analogue of the video tokenizer’s frame-versus-resolution trade,…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.