← All parts of this equation

Equation 9 · Part 5 · Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared

Symbol z_t

Lseq=−∑t=1Tlog⁡pθ ⁣(zt∣z<t),zt∈{0,1,…,V−1}\mathcal{L}_{\text{seq}} = -\sum_{t=1}^{T} \log p_\theta\!\left(z_t \mid z_{<t}\right), \qquad z_t \in \{0, 1, \dots, V-1\}
ztz_t

What this part means

allowed to denote.

Its job in the formula

ztz_t is one of the signed contributions combined to compute the quantity on the left.

Where the article explains it

The same cross-entropy objective a text-only language model uses is the entire training objective here — the only thing that changed is what ztz_t is allowed to denote.

The passage around this formula

…text-to-image generation and entity-linking without task-specific heads [ 8 ] . The same cross-entropy objective a text-only language model uses is the entire training objective here — the only thing that changed is what ztz_t is allowed to denote. Whether index ztz_t names a subword, a 16-by-16 image patch code, or a discretised robot joint angle is invisible to the loss function; the difficulty this section documents is not in the objective but in keeping training numerically stable once the vocabulary spans…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.