← All parts of this equation

Equation 9 · Part 13 · Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared

Ending index or upper bound: T

Lseq=−∑t=1Tlog⁡pθ ⁣(zt∣z<t),zt∈{0,1,…,V−1}\mathcal{L}_{\text{seq}} = -\sum_{t=1}^{T} \log p_\theta\!\left(z_t \mid z_{<t}\right), \qquad z_t \in \{0, 1, \dots, V-1\}
TT

What this part means

This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.

Its job in the formula

T appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

The passage around this formula

Meta’s Chameleon commits to the same principle at a much larger scale and is candid that stability, not capability, was the hard engineering problem. It describes itself as “a family of early-fusion token-based mixed-modal models” that interleaves image and text tokens in one sequence, generating either kind at any position, and its authors state directly that this “requires a stable training approach from inception, an alignment recipe, and an architectural parameterization tailored for the early-fusion, token-based, mixed-modal setting” [ 7 ] . The paper documents the failure mode this addresses concretely: without a query-key normalisation step controlling the growth of attention logits,…

Read this part in the article →

Learn the underlying idea

Σ adds a collection of terms. Π multiplies them. The lower and upper labels tell you which terms belong to the collection.

Open the illustrated sums and products: repeat an operation over an index guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.