Equation 9 · Part 5 · Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared
Symbol z_t
What this part means
allowed to denote.
Its job in the formula
is one of the signed contributions combined to compute the quantity on the left.
Full expression→Symbol z_t→Article meaning
Where the article explains it
The same cross-entropy objective a text-only language model uses is the entire training objective here — the only thing that changed is what is allowed to denote.
The passage around this formula
…text-to-image generation and entity-linking without task-specific heads [ 8 ] . The same cross-entropy objective a text-only language model uses is the entire training objective here — the only thing that changed is what is allowed to denote. Whether index names a subword, a 16-by-16 image patch code, or a discretised robot joint angle is invisible to the loss function; the difficulty this section documents is not in the objective but in keeping training numerically stable once the vocabulary spans…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the surrounding passage
- [7] Chameleon: Mixed-Modal Early-Fusion Foundation Models ↗
- [8] CM3: A Causal Masked Multimodal Model of the Internet ↗
These citations provide research context; check each source for the exact claim it supports.