Equation 8 · How Training Data and Synthetic Data Actually Work
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol n
n is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The word “irreversible” is precise and load-bearing: it does not mean the process is fast, or that any synthetic data at all causes it, or that mixing synthetic data with real data necessarily triggers it. It means that once the tails of the distribution are lost through repeated recursive resampling without fresh real data anchoring each generation, that specific information cannot be recovered from later generations of the same recursive process, because it never encoded a copy of those tails anywhere in the chain. This is the load-bearing distinction between the phi-model result and the collapse result, and conflating them is the most common misreading of both papers. Phi-1’s synthetic…
Read the full surrounding passage
The word “irreversible” is precise and load-bearing: it does not mean the process is fast, or that any synthetic data at all causes it, or that mixing synthetic data with real data necessarily triggers it. It means that once the tails of the distribution are lost through repeated recursive resampling without fresh real data anchoring each generation, that specific information cannot be recovered from later generations of the same recursive process, because it never encoded a copy of those tails anywhere in the chain. This is the load-bearing distinction between the phi-model result and the collapse result, and conflating them is the most common misreading of both papers. Phi-1’s synthetic textbooks are generated once, by a stronger external teacher model (GPT-3.5), and then mixed with real filtered web text as a curriculum for a smaller student model — a single generation of teacher-to-student distillation, not a recursive chain of a model consuming its own prior outputs across many generations [ 4 ] . Model collapse, as demonstrated in the Shumailov paper, is specifically about the recursive case: generation n+1 trained substantially on the sampled output of generation n , repeated across many generations, with the real data anchor diluted or absent [ 1 ] . Curated, teacher-generated synthetic data used once as a supplement to real data is a different mechanism from an uncurated corpus that increasingly consists of unlabeled prior-model output scraped back off the open web across successive training cycles — and the second describes a real and growing risk as more of the public web itself is populated by earlier models’ output, which is the actual mechanism by which recursive contamination could occur without any lab intending it.
Sources cited in the surrounding passage
- [4] Textbooks Are All You Need ↗
- [1] The Curse of Recursion: Training on Generated Data Makes Models Forget ↗
These citations give research context. Read each source to check which claims it supports.
Return to How Training Data and Synthetic Data Actually Work