← All parts of this equation

Equation 8 · Part 2 · Comparing the Main Approaches to Training Data and Synthetic Data

Symbol c_gen

caccepted=cgenpc_{\mathrm{accepted}} = \frac{c_{\mathrm{gen}}}{p}
cgenc_{\mathrm{gen}}

What this part means

cgc_gen occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Its job in the formula

cgc_gen occurs above the fraction bar. The numerator is divided by the entire denominator below it.

The passage around this formula

Rejection sampling generates multiple candidate completions per prompt from a model and keeps only the ones a separate scoring step accepts, using the survivors as supervised fine-tuning data. Llama 2’s RLHF pipeline used exactly this for its first four rounds, generating candidates from the 70-billion-parameter model, scoring them against a trained reward model, and only introducing proximal policy optimisation as a second mechanism in later rounds once the returns from sampling alone began to taper [ 6 ] . A second, independent demonstration in mathematical reasoning makes the mechanism even more explicit: rejection sampling fine-tuning collects correct reasoning paths directly from a…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.