Equation 4 · Training on Your Own Output: Synthetic Data and What It Does to a Distribution
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol X_1
is an input to the expression that computes the quantity on the left.
Symbol X_n
is an input to the expression that computes the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Let be the model fitted at generation k , and the fitting procedure. The classical collapse setting is . in which generation k+1 sees n samples drawn from its immediate predecessor and nothing else . The original data are gone. No filter selects among the samples. The lineage is single. Under those conditions, degradation compounds because there is no channel by which an error introduced at generation k can ever be corrected.
Sources cited in the article section
- [2] Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data ↗
- [1] AI models collapse when trained on recursively generated data ↗
- [5] Position: Model Collapse Does Not Mean What You Think ↗
These citations give research context. Read each source to check which claims it supports.
Return to Training on Your Own Output: Synthetic Data and What It Does to a Distribution