← Mathematical compendium

Published equation contexts

Ldiffusion=Ex0, ϵ∼N(0,I), t[ ∥ϵ−ϵθ(xt,t,c)∥2 ],xt=αˉt x0+1−αˉt ϵ\mathcal{L}_{\text{diffusion}} = \mathbb{E}_{x_0,\, \epsilon \sim \mathcal{N}(0, I),\, t} \Big[\, \big\| \epsilon - \epsilon_\theta(x_t, t, c) \big\|^2 \,\Big], \qquad x_t = \sqrt{\bar{\alpha}_t}\, x_0 + \sqrt{1 - \bar{\alpha}_t}\, \epsilon

Why this formula appears here

The fourth lineage answers a different question than the first three. Adapters, native pretraining and unified tokenization are all, in the end, recipes for understanding or jointly representing multiple modalities; diffusion is a recipe for generating one, typically conditioned on another, and it does not use next-token prediction at all. A diffusion model learns to reverse a fixed process that gradually adds Gaussian noise to data, training a network to predict the noise component at each step: Ldiffusion=Ex0, ϵ∼N(0,I), t[ ∥ϵ−ϵθ(xt,t,c)∥2 ],xt=αˉt x0+1−αˉt ϵ\mathcal{L}_{\text{diffusion}} = \mathbb{E}_{x_0,\, \epsilon \sim \mathcal{N}(0, I),\, t} \Big[\, \big\| \epsilon - \epsilon_\theta(x_t, t, c) \big\|^2 \,\Big], \qquad x_t = \sqrt{\bar{\alpha}_t}\, x_0 + \sqrt{1 - \bar{\alpha}_t}\, \epsilon. where c is a conditioning signal — most often a text embedding — and generation runs the process in reverse, starting from pure noise and repeatedly subtracting a predicted…

Read the full article-specific guide →

Read the representative guide

Ldiffusion\mathcal{L}_{\text{diffusion}}

Symbol L_diffusion

LdL_diffusion is computed from the expected values combined on the right.

Read this term in its guide →
Ex0, ϵ∼N(0,I), t\mathbb{E}_{x_0,\, \epsilon \sim \mathcal{N}(0, I),\, t}

Symbol E_x_0, epsilon sim N(0, I), t

E_x0x_0, epsilon sim N(0, I), t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
ϵ\epsilon

Symbol epsilon

epsilon appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
ϵθ\epsilon_\theta

Symbol epsilon_θ

epsilon_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
xtx_t

Symbol x_t

xtx_t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
tt

Symbol t

t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
cc

Symbol c

a conditioning signal — most often a text embedding — and generation runs the process in reverse, starting from pure noise and repeatedly subtracting a predicted noise estimate.

Read this term in its guide →
αˉt\bar{\alpha}_t

Symbol barα_t

barα_t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
x0x_0

Symbol x_0

x0x_0 appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Ldiffusion=Ex0, ϵ∼N(0,I), t[ ∥ϵ−ϵθ(xt,t,c)∥2 ],xt=αˉt x0+1−αˉt ϵ\mathcal{L}_{\text{diffusion}} = \mathbb{E}_{x_0,\, \epsilon \sim \mathcal{N}(0, I),\, t} \Big[\, \big\| \epsilon - \epsilon_\theta(x_t, t, c) \big\|^2 \,\Big], \qquad x_t = \sqrt{\bar{\alpha}_t}\, x_0 + \sqrt{1 - \bar{\alpha}_t}\, \epsilon

Equation 12 · Foundation Models

Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

The fourth lineage answers a different question than the first three. Adapters, native pretraining and unified tokenization are all, in the end, recipes for understanding or jointly representing multiple modalities; diffusion is a recipe for generating one, typically conditioned on another, and it does not use next-token prediction at all. A diffusion model learns to reverse a fixed process that gradually adds Gaussian noise to data, training a network to predict the noise component at each step: Ldiffusion=Ex0, ϵ∼N(0,I), t[ ∥ϵ−ϵθ(xt,t,c)∥2 ],xt=αˉt x0+1−αˉt ϵ\mathcal{L}_{\text{diffusion}} = \mathbb{E}_{x_0,\, \epsilon \sim \mathcal{N}(0, I),\, t} \Big[\, \big\| \epsilon - \epsilon_\theta(x_t, t, c) \big\|^2 \,\Big], \qquad x_t = \sqrt{\bar{\alpha}_t}\, x_0 + \sqrt{1 - \bar{\alpha}_t}\, \epsilon. where c is a conditioning signal — most often a text embedding — and generation runs the process in reverse, starting from pure noise and repeatedly subtracting a predicted…

Meanings in this article

  • cc: a conditioning signal — most often a text embedding — and generation runs the process in reverse, starting from pure noise and repeatedly subtracting a predicted noise estimate.
Equation guide → · Article →