← Back to article

Equation 7 · How Training Data and Synthetic Data Actually Work

What does this equation mean?

n+1n+1

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

nn

Symbol n

n+1 means one index position after n. In the article, use the nearby sentence to see what each position counts.

Understand this part →

addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

The word “irreversible” is precise and load-bearing: it does not mean the process is fast, or that any synthetic data at all causes it, or that mixing synthetic data with real data necessarily triggers it. It means that once the tails of the distribution are lost through repeated recursive resampling without fresh real data anchoring each generation, that specific information cannot be recovered from later generations of the same recursive process, because it never encoded a copy of those tails anywhere in the chain. This is the load-bearing distinction between the phi-model result and the collapse result, and conflating them is the most common misreading of both papers. Phi-1’s synthetic…
Read the full surrounding passage
The word “irreversible” is precise and load-bearing: it does not mean the process is fast, or that any synthetic data at all causes it, or that mixing synthetic data with real data necessarily triggers it. It means that once the tails of the distribution are lost through repeated recursive resampling without fresh real data anchoring each generation, that specific information cannot be recovered from later generations of the same recursive process, because it never encoded a copy of those tails anywhere in the chain. This is the load-bearing distinction between the phi-model result and the collapse result, and conflating them is the most common misreading of both papers. Phi-1’s synthetic textbooks are generated once, by a stronger external teacher model (GPT-3.5), and then mixed with real filtered web text as a curriculum for a smaller student model — a single generation of teacher-to-student distillation, not a recursive chain of a model consuming its own prior outputs across many generations [ 4 ] . Model collapse, as demonstrated in the Shumailov paper, is specifically about the recursive case: generation n+1 trained substantially on the sampled output of generation n , repeated across many generations, with the real data anchor diluted or absent [ 1 ] . Curated, teacher-generated synthetic data used once as a supplement to real data is a different mechanism from an uncurated corpus that increasingly consists of unlabeled prior-model output scraped back off the open web across successive training cycles — and the second describes a real and growing risk as more of the public web itself is populated by earlier models’ output, which is the actual mechanism by which recursive contamination could occur without any lab intending it.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to How Training Data and Synthetic Data Actually Work

See this formula across 1 published context →

Browse the mathematical compendium →