← All parts of this equation

Equation 12 · Part 2 · How Multimodal Models Actually Handle Video, Audio, and Space

Symbol f_s

R=fsh⋅nq⋅log⁡2K bits per second.R = \frac{f_s}{h} \cdot n_q \cdot \log_2 K \text{ bits per second.}
fsf_s

What this part means

fsf_s occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Its job in the formula

fsf_s occurs above the fraction bar. The numerator is divided by the entire denominator below it.

The passage around this formula

…encoder-decoder pair with a residual vector quantizer, trained jointly and shown to outperform prior codecs at comparable bitrates [ 7 ] . The mechanism gives a concrete, computable token rate. For a signal sampled at fsf_s hertz, encoded with hop length h samples per frame, and quantized with nqn_q residual stages of codebook size K each, the bitrate is R=fsh⋅nq⋅log⁡2K bits per second.R = \frac{f_s}{h} \cdot n_q \cdot \log_2 K \text{ bits per second.}. Every added quantizer stage buys reconstruction fidelity at a fixed, computable token-rate cost — the audio-codec analogue of the video…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.