← All parts of this equation

Equation 8 · Part 2 · Training on Your Own Output: Synthetic Data and What It Does to a Distribution

Symbol F

p^k+1=F(D0∪S1∪⋯∪Sk),\hat{p}_{k+1} = \mathcal{F}\left( D_0 \cup S_1 \cup \cdots \cup S_k \right),
F\mathcal{F}

What this part means

the fitting procedure.

Its job in the formula

F is an input to the expression that computes the quantity on the left.

Where the article explains it

Let p^k\hat{p}_k be the model fitted at generation k , and F\mathcal{F} the fitting procedure.

The passage around this formula

Change one term and the analysis changes with it. Consider instead p^k+1=F(D0∪S1∪⋯∪Sk)\hat{p}_{k+1} = \mathcal{F}\left( D_0 \cup S_1 \cup \cdots \cup S_k \right). where D0D_0 is the original real corpus and SjS_j the synthetic output of generation j . Gerstgrasser and colleagues make exactly this substitution and report that while replacing real data with each generation’s synthetic data does tend toward collapse, accumulating successive generations alongside the original real data avoids it — across transformers, diffusion models and variational autoencoders — and they prove that under accumulation the test error has a finite upper bound independent of the number of iterations [ 2 ] . The accumulating case is also the more accurate description of the actual web, which…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.