Equation 12 · Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol L_diffusion
iffusion is computed from the expected values combined on the right.
Symbol E_x_0, epsilon sim N(0, I), t
E_, epsilon sim N(0, I), t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol epsilon
epsilon appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol epsilon_θ
epsilon_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol x_t
appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol t
t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol c
a conditioning signal — most often a text embedding — and generation runs the process in reverse, starting from pure noise and repeatedly subtracting a predicted noise estimate.
Symbol barα_t
barα_t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol x_0
appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The fourth lineage answers a different question than the first three. Adapters, native pretraining and unified tokenization are all, in the end, recipes for understanding or jointly representing multiple modalities; diffusion is a recipe for generating one, typically conditioned on another, and it does not use next-token prediction at all. A diffusion model learns to reverse a fixed process that gradually adds Gaussian noise to data, training a network to predict the noise component at each step: . where c is a conditioning signal — most often a text embedding — and generation runs the process in reverse, starting from pure noise and repeatedly subtracting a predicted…
Read the full surrounding passage
The fourth lineage answers a different question than the first three. Adapters, native pretraining and unified tokenization are all, in the end, recipes for understanding or jointly representing multiple modalities; diffusion is a recipe for generating one, typically conditioned on another, and it does not use next-token prediction at all. A diffusion model learns to reverse a fixed process that gradually adds Gaussian noise to data, training a network to predict the noise component at each step: . where c is a conditioning signal — most often a text embedding — and generation runs the process in reverse, starting from pure noise and repeatedly subtracting a predicted noise estimate.
Sources cited in the article section
- [9] High-Resolution Image Synthesis with Latent Diffusion Models ↗
- [10] Hierarchical Text-Conditional Image Generation with CLIP Latents ↗
- [11] Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding ↗
- [12] Any-to-Any Generation via Composable Diffusion ↗
These citations give research context. Read each source to check which claims it supports.