← All parts of this equation

Equation 8 · Part 4 · Training on Your Own Output: Synthetic Data and What It Does to a Distribution

Symbol S_1

p^k+1=F(D0∪S1∪⋯∪Sk),\hat{p}_{k+1} = \mathcal{F}\left( D_0 \cup S_1 \cup \cdots \cup S_k \right),
S1S_1

What this part means

S1S_1 is an input to the expression that computes the quantity on the left.

Its job in the formula

S1S_1 is an input to the expression that computes the quantity on the left.

The passage around this formula

Change one term and the analysis changes with it. Consider instead p^k+1=F(D0∪S1∪⋯∪Sk)\hat{p}_{k+1} = \mathcal{F}\left( D_0 \cup S_1 \cup \cdots \cup S_k \right). where D0D_0 is the original real corpus and SjS_j the synthetic output of generation j . Gerstgrasser and colleagues make exactly this substitution and report that while replacing real data with each generation’s synthetic data does tend toward collapse, accumulating successive generations alongside the original real data avoids it — across transformers, diffusion models and variational autoencoders — and they prove that under accumulation the test error has a finite upper bound independent of the number of iterations [ 2 ] . The accumulating case is also the more accurate description of the actual web, which…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.