← Back to article

Equation 8 · Comparing the Main Approaches to Training Data and Synthetic Data

What does this equation mean?

caccepted=cgenpc_{\mathrm{accepted}} = \frac{c_{\mathrm{gen}}}{p}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Start withc_gen
Divide byp
This relates toc_accepted
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

cacceptedc_{\mathrm{accepted}}

Symbol c_accepted

cac_accepted is part of the quantity the equation computes from the expression on the right.

Understand this part →

cgenc_{\mathrm{gen}}

Symbol c_gen

cgc_gen occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

pp

Symbol p

the fraction.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Rejection sampling generates multiple candidate completions per prompt from a model and keeps only the ones a separate scoring step accepts, using the survivors as supervised fine-tuning data. Llama 2’s RLHF pipeline used exactly this for its first four rounds, generating candidates from the 70-billion-parameter model, scoring them against a trained reward model, and only introducing proximal policy optimisation as a second mechanism in later rounds once the returns from sampling alone began to taper [ 6 ] . A second, independent demonstration in mathematical reasoning makes the mechanism even more explicit: rejection sampling fine-tuning collects correct reasoning paths directly from a…
Read the full surrounding passage
Rejection sampling generates multiple candidate completions per prompt from a model and keeps only the ones a separate scoring step accepts, using the survivors as supervised fine-tuning data. Llama 2’s RLHF pipeline used exactly this for its first four rounds, generating candidates from the 70-billion-parameter model, scoring them against a trained reward model, and only introducing proximal policy optimisation as a second mechanism in later rounds once the returns from sampling alone began to taper [ 6 ] . A second, independent demonstration in mathematical reasoning makes the mechanism even more explicit: rejection sampling fine-tuning collects correct reasoning paths directly from a supervised model’s own sampled outputs, with no additional human annotation, and combining rejection samples pooled from multiple models raised a 7-billion-parameter LLaMA’s GSM8K accuracy from a 35.9 percent supervised baseline to 49.3 percent [ 7 ] . The method’s economics follow directly from its mechanism: if a checker accepts a fraction p of generated candidates, the expected cost of producing one accepted sample scales as caccepted=cgenpc_{\mathrm{accepted}} = \frac{c_{\mathrm{gen}}}{p}. so as a task gets harder and p falls, rejection sampling gets steadily more expensive for the same yield — precisely the pattern that shows up in Llama 2’s own pipeline as a second mechanism being layered on once sampling alone stopped paying for itself.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Comparing the Main Approaches to Training Data and Synthetic Data

See this formula across 1 published context →

Browse the mathematical compendium →