Equation 12 · Training on Your Own Output: Synthetic Data and What It Does to a Distribution
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the not. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Three structural features explain why. The teacher is fixed : nothing the student produces flows back into it, so there is no feedback loop to compound. The chain has length one , not k . And the objective is transfer of a known-good distribution, not improvement beyond it, so the success criterion is agreement rather than novelty. The relevant failure mode is entirely different from collapse: a student inherits the teacher’s errors and cannot exceed the teacher’s ceiling on the distilled behaviour.
Sources cited in the article section
- [13] Distilling the Knowledge in a Neural Network ↗
- [14] DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning ↗
These citations give research context. Read each source to check which claims it supports.
Return to Training on Your Own Output: Synthetic Data and What It Does to a Distribution