← All parts of this equation

Equation 12 · Part 9 · How Multimodal Models Actually Handle Video, Audio, and Space

subscript

R=fsh⋅nq⋅log⁡2K bits per second.R = \frac{f_s}{h} \cdot n_q \cdot \log_2 K \text{ bits per second.}
subscript

What this part means

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Its job in the formula

A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.

The passage around this formula

The mechanism that does this at scale is a learned neural codec built around residual vector quantization: an encoder compresses the waveform into a sequence of continuous frames, and each frame is then quantized in successive stages, each stage encoding what the previous stage’s codebook missed. SoundStream is the reference architecture, described by its authors as an end-to-end neural audio codec built on a convolutional encoder-decoder pair with a residual vector quantizer, trained jointly and shown to outperform prior codecs at comparable bitrates [ 7 ] . The mechanism gives a concrete, computable token rate. For a signal sampled at fsf_s hertz, encoded with hop length h samples per frame,…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.