Equation 10 · Part 1 · Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared
Symbol z_t
What this part means
allowed to denote.
Its job in the formula
is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol z_t→Article meaning
Where the article explains it
The same cross-entropy objective a text-only language model uses is the entire training objective here — the only thing that changed is what is allowed to denote.
The passage around this formula
The same cross-entropy objective a text-only language model uses is the entire training objective here — the only thing that changed is what is allowed to denote. Whether index names a subword, a 16-by-16 image patch code, or a discretised robot joint angle is invisible to the loss function; the difficulty this section documents is not in the objective but in keeping training numerically stable once the vocabulary spans quantities of such different statistical character.
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the article section
- [6] A Generalist Agent ↗
- [7] Chameleon: Mixed-Modal Early-Fusion Foundation Models ↗
- [8] CM3: A Causal Masked Multimodal Model of the Internet ↗
These citations provide research context; check each source for the exact claim it supports.