← Back to article

Equation 8 · Training on Your Own Output: Synthetic Data and What It Does to a Distribution

What does this equation mean?

p^k+1=F(D0∪S1∪⋯∪Sk),\hat{p}_{k+1} = \mathcal{F}\left( D_0 \cup S_1 \cup \cdots \cup S_k \right),

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationsF( D_0 cup S_1 cup × s cup S_k )
Result or conditionhatp_k+1
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

p^k+1\hat{p}_{k+1}

Symbol hatp_k+1

hatpkp_k+1 is part of the quantity the equation computes from the expression on the right.

Understand this part →

F\mathcal{F}

Symbol F

the fitting procedure.

Understand this part →

D0D_0

Symbol D_0

the original real corpus and SjS_j the synthetic output of generation j.

Understand this part →

S1S_1

Symbol S_1

S1S_1 is an input to the expression that computes the quantity on the left.

Understand this part →

SkS_k

Symbol S_k

SkS_k is an input to the expression that computes the quantity on the left.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Change one term and the analysis changes with it. Consider instead p^k+1=F(D0∪S1∪⋯∪Sk)\hat{p}_{k+1} = \mathcal{F}\left( D_0 \cup S_1 \cup \cdots \cup S_k \right). where D0D_0 is the original real corpus and SjS_j the synthetic output of generation j . Gerstgrasser and colleagues make exactly this substitution and report that while replacing real data with each generation’s synthetic data does tend toward collapse, accumulating successive generations alongside the original real data avoids it — across transformers, diffusion models and variational autoencoders — and they prove that under accumulation the test error has a finite upper bound independent of the number of iterations [ 2 ] . The accumulating case is also the more accurate description of the actual web, which…
Read the full surrounding passage
Change one term and the analysis changes with it. Consider instead p^k+1=F(D0∪S1∪⋯∪Sk)\hat{p}_{k+1} = \mathcal{F}\left( D_0 \cup S_1 \cup \cdots \cup S_k \right). where D0D_0 is the original real corpus and SjS_j the synthetic output of generation j . Gerstgrasser and colleagues make exactly this substitution and report that while replacing real data with each generation’s synthetic data does tend toward collapse, accumulating successive generations alongside the original real data avoids it — across transformers, diffusion models and variational autoencoders — and they prove that under accumulation the test error has a finite upper bound independent of the number of iterations [ 2 ] . The accumulating case is also the more accurate description of the actual web, which does not delete last year’s pages when this year’s are published.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Training on Your Own Output: Synthetic Data and What It Does to a Distribution

See this formula across 1 published context →

Browse the mathematical compendium →