← Back to article

Equation 12 · Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared

What does this equation mean?

Ldiffusion=Ex0, ϵ∼N(0,I), t[ ∥ϵ−ϵθ(xt,t,c)∥2 ],xt=αˉt x0+1−αˉt ϵ\mathcal{L}_{\text{diffusion}} = \mathbb{E}_{x_0,\, \epsilon \sim \mathcal{N}(0, I),\, t} \Big[\, \big\| \epsilon - \epsilon_\theta(x_t, t, c) \big\|^2 \,\Big], \qquad x_t = \sqrt{\bar{\alpha}_t}\, x_0 + \sqrt{1 - \bar{\alpha}_t}\, \epsilon

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

Ldiffusion\mathcal{L}_{\text{diffusion}}

Symbol L_diffusion

LdL_diffusion is computed from the expected values combined on the right.

Understand this part →

Ex0, ϵ∼N(0,I), t\mathbb{E}_{x_0,\, \epsilon \sim \mathcal{N}(0, I),\, t}

Symbol E_x_0, epsilon sim N(0, I), t

E_x0x_0, epsilon sim N(0, I), t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

ϵ\epsilon

Symbol epsilon

epsilon appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

ϵθ\epsilon_\theta

Symbol epsilon_θ

epsilon_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

xtx_t

Symbol x_t

xtx_t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

tt

Symbol t

t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

cc

Symbol c

a conditioning signal — most often a text embedding — and generation runs the process in reverse, starting from pure noise and repeatedly subtracting a predicted noise estimate.

Understand this part →

αˉt\bar{\alpha}_t

Symbol barα_t

barα_t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

x0x_0

Symbol x_0

x0x_0 appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
√

√

Take a square root.

Understand this part →

addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subtraction

subtraction

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

The fourth lineage answers a different question than the first three. Adapters, native pretraining and unified tokenization are all, in the end, recipes for understanding or jointly representing multiple modalities; diffusion is a recipe for generating one, typically conditioned on another, and it does not use next-token prediction at all. A diffusion model learns to reverse a fixed process that gradually adds Gaussian noise to data, training a network to predict the noise component at each step: Ldiffusion=Ex0, ϵ∼N(0,I), t[ ∥ϵ−ϵθ(xt,t,c)∥2 ],xt=αˉt x0+1−αˉt ϵ\mathcal{L}_{\text{diffusion}} = \mathbb{E}_{x_0,\, \epsilon \sim \mathcal{N}(0, I),\, t} \Big[\, \big\| \epsilon - \epsilon_\theta(x_t, t, c) \big\|^2 \,\Big], \qquad x_t = \sqrt{\bar{\alpha}_t}\, x_0 + \sqrt{1 - \bar{\alpha}_t}\, \epsilon. where c is a conditioning signal — most often a text embedding — and generation runs the process in reverse, starting from pure noise and repeatedly subtracting a predicted…
Read the full surrounding passage
The fourth lineage answers a different question than the first three. Adapters, native pretraining and unified tokenization are all, in the end, recipes for understanding or jointly representing multiple modalities; diffusion is a recipe for generating one, typically conditioned on another, and it does not use next-token prediction at all. A diffusion model learns to reverse a fixed process that gradually adds Gaussian noise to data, training a network to predict the noise component at each step: Ldiffusion=Ex0, ϵ∼N(0,I), t[ ∥ϵ−ϵθ(xt,t,c)∥2 ],xt=αˉt x0+1−αˉt ϵ\mathcal{L}_{\text{diffusion}} = \mathbb{E}_{x_0,\, \epsilon \sim \mathcal{N}(0, I),\, t} \Big[\, \big\| \epsilon - \epsilon_\theta(x_t, t, c) \big\|^2 \,\Big], \qquad x_t = \sqrt{\bar{\alpha}_t}\, x_0 + \sqrt{1 - \bar{\alpha}_t}\, \epsilon. where c is a conditioning signal — most often a text embedding — and generation runs the process in reverse, starting from pure noise and repeatedly subtracting a predicted noise estimate.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared

See this formula across 1 published context →

Browse the mathematical compendium →