← All parts of this equation

Equation 12 · Part 4 · How Multimodal Models Actually Handle Video, Audio, and Space

Symbol n_q

R=fsh⋅nq⋅log⁡2K bits per second.R = \frac{f_s}{h} \cdot n_q \cdot \log_2 K \text{ bits per second.}
nqn_q

What this part means

nqn_q is one factor in the product that computes the quantity on the left.

Its job in the formula

nqn_q is one factor in the product that computes the quantity on the left.

The passage around this formula

…shown to outperform prior codecs at comparable bitrates [ 7 ] . The mechanism gives a concrete, computable token rate. For a signal sampled at fsf_s hertz, encoded with hop length h samples per frame, and quantized with nqn_q residual stages of codebook size K each, the bitrate is R=fsh⋅nq⋅log⁡2K bits per second.R = \frac{f_s}{h} \cdot n_q \cdot \log_2 K \text{ bits per second.}. Every added quantizer stage buys reconstruction fidelity at a fixed, computable token-rate cost — the audio-codec analogue of the video tokenizer’s frame-versus-resolution trade, and just as unavoidable.

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.