← All parts of this equation

Equation 9 · Part 6 · Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared

Symbol z_<t

Lseq=−∑t=1Tlog⁡pθ ⁣(zt∣z<t),zt∈{0,1,…,V−1}\mathcal{L}_{\text{seq}} = -\sum_{t=1}^{T} \log p_\theta\!\left(z_t \mid z_{<t}\right), \qquad z_t \in \{0, 1, \dots, V-1\}
z<tz_{<t}

What this part means

z_<t is one of the signed contributions combined to compute the quantity on the left.

Its job in the formula

z_<t is one of the signed contributions combined to compute the quantity on the left.

The passage around this formula

Meta’s Chameleon commits to the same principle at a much larger scale and is candid that stability, not capability, was the hard engineering problem. It describes itself as “a family of early-fusion token-based mixed-modal models” that interleaves image and text tokens in one sequence, generating either kind at any position, and its authors state directly that this “requires a stable training approach from inception, an alignment recipe, and an architectural parameterization tailored for the early-fusion, token-based, mixed-modal setting” [ 7 ] . The paper documents the failure mode this addresses concretely: without a query-key normalisation step controlling the growth of attention logits,…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.