Five ways to get training data, and a rule about ranking them
“Synthetic data” is usually spoken of as though it names one pipeline. It does not. At least five structurally distinct methods now supply the tokens a model trains on, and they share almost nothing except that a machine, not a printing press, produced the text. Distillation copies the output distribution of a larger, already-trained teacher into a smaller student. Self-improvement, sometimes called self-play, has a model generate data from its own outputs and keeps only what a scoring or verification step lets through. Formally verified generation produces data in domains — mathematics, formal proof, code — where an external checker with no learned parameters at all can confirm correctness. Human-annotated and curated data is written or ranked by paid annotators against instructions. Web-scale scraped-and-filtered data is harvested from the open internet and cut down by deduplication and quality classifiers. A sixth question sits beside the first five and is addressed on its own terms later: once a base model exists, what kind of data should shape it further — data selected by rejection sampling, preference data produced by human or AI feedback, or text a stronger model writes directly?
This article’s job is to compare these approaches structurally: what each one actually requires, what has been demonstrated with it and at what scale, and what it cannot do. It will not produce a ranking across them, and that omission is deliberate rather than an oversight. The results cited below come from different model families, different parameter counts, different tasks, and largely undisclosed and certainly unmatched compute budgets. A number describing how well distillation worked on a 1.3-billion-parameter code model has no valid comparison to a number describing how well formal verification worked on a system that is not even a general-purpose language model. Where a head-to-head is not possible, this piece says so explicitly rather than manufacturing one, in keeping with the discipline that governs comparative claims in this publication generally: no cross-approach ranking gets built from conditions that were never controlled to allow one.
Distillation: transferring a fixed, known-good distribution
Distillation is the oldest of the five and the easiest to state precisely. A student model is trained not on hard labels but on a teacher’s softened output distribution, on the reasoning that the relative probabilities the teacher assigns to wrong answers carry information a one-hot label discards entirely [1]. Where logits are unavailable — the ordinary case when the teacher is a closed API — the same idea is applied at one remove: the teacher’s sampled text stands in for its distribution, and the student is trained on that text directly. Writing
The point of writing it out is the structure it exposes, not the arithmetic: the teacher term is fixed throughout training, so nothing the student produces ever feeds back into it. There is no loop to compound and no possibility of the drift that shows up when a model is fitted repeatedly to its own unfiltered output.
What distillation has actually demonstrated, and at what scale, is now well documented. The peer-reviewed account of DeepSeek-R1 describes distilling the reasoning behaviour of a 671-billion-parameter reinforcement-learned model into six dense checkpoints ranging from 1.5 billion to 70 billion parameters, using roughly 800,000 reasoning samples generated by the large model, with no further reinforcement-learning stage applied to the students — pure supervised fine-tuning on a fixed teacher’s output was sufficient to produce distilled models the paper reports as exceeding their own base checkpoints’ prior reasoning performance [2]. A second, quieter demonstration sits inside Llama 2’s post-training pipeline: rejection sampling, described in the next section, was applied only to the largest 70-billion-parameter chat model, and every smaller variant was fine-tuned on samples drawn from that 70-billion-parameter output — the paper’s own authors describe this explicitly as distilling the large model’s capabilities into the smaller ones [6]. Distillation, in other words, is not confined to a separate stage; it is routinely the last step applied to whatever a more expensive method already produced.
The demonstrated failure mode is equally specific and worth stating precisely rather than left implicit: a student trained this way inherits the teacher’s errors on the distilled behaviour and has no documented mechanism for exceeding the teacher’s ceiling on that behaviour. Nothing in the cited results shows a distilled student surpassing its teacher on the capability being transferred; both results above describe transfer, not improvement beyond the source.
Self-improvement and self-play: the model marks its own homework, if something else grades it
Self-improvement is the case that most resembles the closed-loop recursion documented elsewhere as producing collapse — a model’s own output becomes training data — with one load-bearing difference: an external step decides what survives. Where nothing sorts the output, the failure is real and well characterised: recursive fitting on unfiltered model-generated text has been shown, in controlled experiments retraining OPT-125m across generations, to erase low-probability tail behaviour first and then collapse the retained distribution’s variance, with reported perplexity degradations of roughly 20 to 28 points when no original data survived across generations [14]. Every result described in this section depends on breaking that loop with a step the generating pass does not control.
Self-Instruct is the clearest early demonstration. A model bootstraps 52,000 instruction–input–output examples from 175 seed tasks by repeatedly prompting itself, and invalid or near-duplicate generations are filtered out before any fine-tuning happens — the filter, not the generation step, decides what enters the training set. Applied to vanilla GPT-3, the resulting fine-tune achieved a 33 percentage-point absolute improvement on Super-NaturalInstructions and matched the performance level of InstructGPT’s first release, leaving only a small remaining gap on expert-curated novel tasks [10]. Nothing here required a stronger external teacher; the same weights generated, and a separate filtering pass selected.
The reinforcement-learning version of self-improvement replaces a discrete filter with a continuous reward, and DeepSeek-R1-Zero is the clearest published account of it: a base model trained by reinforcement learning directly against a verifiable correctness signal on math and code tasks, without any supervised fine-tuning cold start and without human-annotated reasoning traces, produced self-verification and strategy-revision behaviour that the authors describe as emergent, and raised AIME 2024 pass@1 from 15.6 percent to 77.9 percent over training [2]. The self-improvement loop here is sound for the same structural reason distillation is sound: the reward is checkable, so it functions as the outside step, even though it arrives from a program rather than a second model.
A harder case is a model checking itself with no external checker at all. Self-Rewarding Language Models fine-tuned a 70-billion-parameter Llama 2 through three rounds of iterative preference optimisation in which the model being trained also acts as its own judge, using LLM-as-a-judge prompting to score its own candidate responses and generate the preference pairs used in the next round. After three iterations the resulting model’s win rate on the AlpacaEval 2.0 leaderboard exceeded that of Claude 2, Gemini Pro, and GPT-4 0613 as reported by the same evaluation [3]. This result is real and it is also the weakest-checked case in this section, and the honest caveat deserves stating plainly: the judge is drawn from the same weights being improved, so a systematic blind spot the model has about its own outputs is exactly the kind of error this design has no independent mechanism for catching. That is a structural property of the method, not a criticism of the specific paper, and it is the reason self-improvement sits on a spectrum rather than in a single box — a spectrum this article returns to directly in the section on post-training data sources.
Formally verified generation: where a checker outside the model ends the argument
A narrower category has the strongest guarantee of the five, precisely because its checker is not a model at all. AlphaGeometry trained a language model from scratch on 100 million synthesised geometry theorems and proofs, generated and confirmed correct by a symbolic deduction engine with no learned parameters, to guide that same engine’s search. On a held-out set of 30 recent olympiad-level problems it solved 25 within the standard time limit, against 10 for the previous best method, and the paper reports performance approaching that of an average International Mathematical Olympiad gold medallist [4]. AlphaProof extends the same logic to general mathematics through the Lean proof assistant: a reinforcement-learning agent was trained against roughly 80 million formal problems auto-formalised from natural-language statements, with Lean’s kernel as the sole arbiter of whether a proof counted, and at the 2024 IMO it solved three of five non-geometry problems — including a problem only a handful of human competitors solved — while the combined system with AlphaGeometry 2 reached a score in the silver-medal range [5].
The guarantee in both cases is real and it is also narrower than it first appears, for a reason worth stating precisely rather than glossed over. The checker confirms that a specific formal statement was proved; it says nothing about whether that formal statement faithfully renders the informal problem it was translated from, and the translation step in AlphaProof’s own pipeline used a fine-tuned language model, which is not itself checked by Lean [5]. The verified region ends exactly at that translation boundary. And the boundary matters more broadly than as a footnote to two results: essentially no other domain that models are used for — prose, dialogue, legal reasoning, most of code in practice — has an analogous checker. This is why formally verified generation, whatever its strength within mathematics and formal proof, does not generalise into a solution for the other four approaches’ problems; it solves a different, narrower problem completely, rather than solving the general problem partially.
Human-annotated and curated data: the expensive baseline everything else is measured against
Human annotation is structurally distinct from the previous three because its checker is a person’s judgement rather than a filter, a reward function, or a symbolic engine, which is precisely what lets it supervise domains none of the other approaches can reach — open-ended helpfulness, tone, and safety trade-offs with no formal test and, in the case of a genuinely novel capability, no stronger existing model to copy from. InstructGPT is the canonical demonstration of what that buys: supervised fine-tuning on human-written demonstrations, a reward model fitted to human pairwise comparisons, and reinforcement learning against that reward model produced a 1.3-billion-parameter model whose outputs human raters preferred to the 175-billion-parameter base GPT-3, despite having roughly a hundredth of the parameters [8].
The scale this requires in practice is the more striking fact, and Llama 2’s published pipeline gives an unusually concrete figure for it: the total human preference dataset assembled from all sources totalled roughly 2.9 million binary comparisons, of which about 1.4 million were newly collected by Meta specifically for helpfulness and safety, rated on a four-point preference-strength scale and refreshed continuously across five successive rounds of reinforcement learning so the reward model would not go stale relative to the policy it was scoring [6]. That figure is worth holding next to Self-Instruct’s 52,000 bootstrapped examples or AlphaGeometry’s fully automated 100 million theorems: human annotation does not scale the way generation-based methods do, and every result in this section reflects a lab choosing to spend that cost on a targeted, high-leverage slice of the training signal rather than on its full token budget. It functions, in effect, as the calibration reference the cheaper approaches are checked against — Self-Instruct explicitly reports itself as closing a gap toward InstructGPT rather than surpassing it, and Self-Rewarding Language Models is evaluated against models whose own post-training relied in part on exactly this kind of human-labelled data.
Web-scale scraped-and-filtered data: abundance with a shrinking legal margin
The fifth approach supplies volume none of the others can match, and its demonstrated results are as much about filtering as about scale. FineWeb assembles 15 trillion tokens from 96 Common Crawl snapshots with a documented, ablated deduplication and filtering pipeline; a further-filtered 1.3-trillion-token educational subset, FineWeb-Edu, produces markedly better downstream performance on knowledge- and reasoning-heavy benchmarks such as MMLU and ARC than the less-filtered base collection at matched model size [11]. That comparison is one of the few in this article that is genuinely apples-to-apples, because it holds model architecture and token budget fixed and varies only the filter — and what it shows is that at pretraining scale, the filtering step plays the same structural role deduplication and quality classification play for web data that a verifier plays for self-improvement or a reward model plays for RLHF: it is the channel from outside the raw material that decides what survives.
What makes this approach distinct among the five is a constraint none of the others carry at comparable scale: its legal supply is measurably shrinking. A longitudinal audit of roughly 14,000 web domains underlying corpora such as C4 found that within about a year, robots.txt restrictions newly blocked more than 5 percent of all C4 tokens and over 28 percent of the most actively maintained, highest-quality sources within it, and that once terms-of-service crawling restrictions are counted alongside robots.txt, close to 45 percent of C4 is now restricted in some form [12]. Distillation licenses one model’s outputs; formally verified generation needs no external text beyond problem statements; human annotation is commissioned outright; self-improvement recycles a model’s own already-licensed generations. Web-scale scraping is the one approach among the five whose raw material is contested by parties outside the lab doing the training, at a scale and a rate none of the other four’s supply constraints resemble.
Why these five cannot be honestly ranked against one another
Collect the numbers above in one place and the incomparability becomes concrete rather than abstract. AlphaGeometry’s 25 of 30 olympiad problems, FineWeb-Edu’s MMLU gains over unfiltered FineWeb, Self-Rewarding’s AlpacaEval win rate, and InstructGPT’s human preference rate are four different measurements on four different tasks, produced by systems that do not share an architecture, a parameter count, or a disclosed compute budget. AlphaGeometry is not a general-purpose language model at all; it couples a language model to a symbolic deduction engine that has no analogue in the other four approaches. InstructGPT’s headline comparison is between 1.3 billion and 175 billion parameters; Self-Rewarding’s is a 70-billion-parameter model measured against other labs’ undisclosed frontier systems; FineWeb’s ablations hold model size fixed within the paper but at a scale unrelated to any of the others. None of these differences is a flaw in the individual results. They are simply four separate experiments, and stacking their headline numbers into a single ordered list would manufacture a ranking the underlying evidence does not support.
What can be said honestly is narrower and still useful. Within a single paper’s own controlled ablation — FineWeb versus FineWeb-Edu at matched model size and token budget, or Llama 2’s own comparison across its five successive RLHF rounds at matched model size — a comparison is real, because scale and architecture were actually held fixed. Across papers, the defensible claim is qualitative: this approach has been demonstrated to do this, at this scale, checked by this. Several of the approaches, moreover, are not competitors a lab must choose between so much as stages a single pipeline composes. Llama 2 chains rejection sampling into distillation into RLHF inside one release. DeepSeek-R1 chains reinforcement-learning-driven self-improvement into distillation into six separate model sizes. A framing that asks whether distillation “beats” rejection sampling misdescribes what the labs publishing these results actually did, which was use several of them together.
Post-training data sources: rejection sampling, RLHF and RLAIF preference data, and direct synthetic generation
A narrower version of the same comparison concerns not how a base model’s pretraining corpus is built but how it is shaped afterward, and three distinct sources recur in published post-training pipelines.
Rejection sampling generates multiple candidate completions per prompt from a model and keeps only the ones a separate scoring step accepts, using the survivors as supervised fine-tuning data. Llama 2’s RLHF pipeline used exactly this for its first four rounds, generating candidates from the 70-billion-parameter model, scoring them against a trained reward model, and only introducing proximal policy optimisation as a second mechanism in later rounds once the returns from sampling alone began to taper [6]. A second, independent demonstration in mathematical reasoning makes the mechanism even more explicit: rejection sampling fine-tuning collects correct reasoning paths directly from a supervised model’s own sampled outputs, with no additional human annotation, and combining rejection samples pooled from multiple models raised a 7-billion-parameter LLaMA’s GSM8K accuracy from a 35.9 percent supervised baseline to 49.3 percent [7]. The method’s economics follow directly from its mechanism: if a checker accepts a fraction
so as a task gets harder and
RLHF and RLAIF differ in exactly one place: who produces the preference label the reward model trains on. In RLHF the label comes from a person comparing two outputs — Ouyang and colleagues’ human labellers, or the roughly 1.4 million comparisons Meta collected specifically for Llama 2’s helpfulness and safety reward models [8, 6]. In RLAIF the label comes from another model applying a written set of principles: Constitutional AI has a separate AI system compare response pairs against a constitution and trains the reward model on that AI-generated preference set, reported to achieve comparable behavioural control “with far fewer human labels” than a fully human-labelled pipeline [9]. Self-Rewarding Language Models sits at the far end of the axis: the judge is not a separate model at all but the very policy being trained, scoring its own candidates through the process described earlier [3]. The three sit on one genuine spectrum — independent human judge, separately trained AI judge, same-model-as-judge — along which cost falls as independence from the policy’s own blind spots also falls. That is a real structural ordering along a named axis, a different and more defensible claim than a performance ranking, because it requires no comparison of incomparable benchmark numbers.
Direct synthetic generation is the remaining source, differing from the first three in producing training examples outright rather than labels or accept/reject decisions. Textbooks Are All You Need trained a 1.3-billion-parameter code model, phi-1, on 6 billion tokens of web-filtered “textbook quality” text alongside 1 billion tokens of synthetic textbooks and exercises generated directly by GPT-3.5, reaching 50.6 percent pass@1 on HumanEval and 55.5 percent on MBPP [13]. Structurally this sits closer to distillation than to self-improvement — the generator is a fixed, external, presumed-stronger model, not the trainee’s own weights — but unlike classical logit-based distillation it uses only the teacher’s sampled text, with no access to its underlying distribution. As with the five main approaches, these three post-training sources were demonstrated on different tasks at different scales — Llama 2’s rejection sampling at 70 billion parameters on open-ended chat, the RFT paper’s at 7 billion on grade-school arithmetic, Constitutional AI’s RLAIF on Anthropic’s then-current model, Self-Rewarding’s same-model judging at 70 billion on general instruction-following, phi-1’s direct generation at 1.3 billion on code alone — so the same discipline applies: a structural map, not a leaderboard.
Choosing an approach by what a task actually has, not by a leaderboard
The comparisons above resolve into a small set of questions worth asking before choosing a method. Is there a stronger model whose outputs you are licensed to train on? Distillation, logit-based or sampled-text, is the demonstrated tool. Does the domain have an automatic, non-learned checker — a compiler, a proof assistant, a solver? Formally verified generation is the one category here where self-improvement carries no correlated-blind-spot risk, because the checker shares no weights with the thing being checked. Is the target open-ended — tone, helpfulness, a safety trade-off — with neither a checker nor a stronger teacher available? Human annotation is the only demonstrated approach that reaches those domains, sized to the highest-leverage slice a budget allows and treated as the reference other methods are validated against, not as a full substitute for them. Does the task need bulk pretraining-scale text? Web-scale scraped-and-filtered data is the only approach demonstrated at that order of magnitude, with the caveat that its legal margin now needs active governance rather than a bigger crawler. Is the model already capable enough that its own critique is informative, and is some correlated-blind-spot risk tolerable? Self-improvement works, provided the accepted output is checked against something the generating pass does not control — the rule every sound example above shares, and exactly what fails in the unfiltered recursive case documented elsewhere.
None of these are exclusive choices, and the published pipelines cited throughout this article do not treat them as such. The practical shape of the field, visible in every result above, is composition — several of these methods layered inside one pipeline, each supplying what the others cannot — rather than a competition with a single winner.
Predictions, with the observations that would falsify them
These are forecasts, clearly separated from the sourced comparison above. Horizon: 16 August 2030.
One. Composed pipelines — some combination of distillation, rejection sampling, and RLHF or RLAIF — will remain the reported norm rather than any single method reported alone, because the cited evidence already shows leading labs composing them rather than choosing one. Disconfirmed if a leading 2030 system card attributes its post-training data entirely to one named method with no combination.
Two. RLAIF-style AI-generated preference data will keep expanding its share of high-volume comparison generation without eliminating human-labelled data, because RLAIF’s own published soundness argument still rests on a human-written set of principles and, typically, a human-labelled anchor set. Disconfirmed if a major lab documents a large-scale post-training run with no human-authored preference or demonstration data anywhere in the pipeline.
Three. Formally verified synthetic data’s share of total training tokens will stay small in absolute terms even as its role in reasoning benchmarks grows, because the checkable-domain property covers a narrow slice of what models are actually used for. Disconfirmed if a frontier lab reports formally verified synthetic data as the majority token source for a general-purpose model not specialised to math or code.
Four. Cross-approach claims that rank distillation, self-improvement, formal verification, human annotation, and web-scraped data against one another in a single ordered list will remain rare in credible peer-reviewed venues, because the confound problem named in this article does not shrink as more papers are published. Disconfirmed if a peer-reviewed venue publishes a matched-compute, matched-architecture head-to-head across three or more of these five approaches producing a clean ranking.
None of these requires a discontinuity. They follow from what the cited evidence already shows: five structurally different methods, each demonstrated on its own terms, increasingly composed rather than chosen between.
What to take away
Five structurally different ways of making training data do five different jobs, verified by five different things: a fixed teacher’s distribution, a filter or reward the generating pass does not control, a symbolic checker outside the model entirely, a person’s judgement, and a quality classifier applied to whatever the crawler found. None is a substitute for what the others verify, and none has been run against the others under conditions that would make a ranking honest — different scales, different architectures, mostly undisclosed compute, and in at least one case not even the same kind of system. The frontier pipelines actually published compose several of these rather than picking a winner among them, and the discipline this article has tried to model throughout is the one worth carrying forward: for any number offered as evidence that one approach works, ask which of the five produced it, at what scale, checked by what — and resist the leaderboard that would turn five separate experiments into one.