Symbol L_diffusion
iffusion is computed from the expected values combined on the right.
Read this term in its guide →Published equation contexts
The fourth lineage answers a different question than the first three. Adapters, native pretraining and unified tokenization are all, in the end, recipes for understanding or jointly representing multiple modalities; diffusion is a recipe for generating one, typically conditioned on another, and it does not use next-token prediction at all. A diffusion model learns to reverse a fixed process that gradually adds Gaussian noise to data, training a network to predict the noise component at each step: . where c is a conditioning signal — most often a text embedding — and generation runs the process in reverse, starting from pure noise and repeatedly subtracting a predicted…
iffusion is computed from the expected values combined on the right.
Read this term in its guide →E_, epsilon sim N(0, I), t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →epsilon appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →epsilon_θ appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →a conditioning signal — most often a text embedding — and generation runs the process in reverse, starting from pure noise and repeatedly subtracting a predicted noise estimate.
Read this term in its guide →barα_t appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 12 · Foundation Models
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
The fourth lineage answers a different question than the first three. Adapters, native pretraining and unified tokenization are all, in the end, recipes for understanding or jointly representing multiple modalities; diffusion is a recipe for generating one, typically conditioned on another, and it does not use next-token prediction at all. A diffusion model learns to reverse a fixed process that gradually adds Gaussian noise to data, training a network to predict the noise component at each step: . where c is a conditioning signal — most often a text embedding — and generation runs the process in reverse, starting from pure noise and repeatedly subtracting a predicted…