← Mathematical compendium

Published equation contexts

caccepted=cgenpc_{\mathrm{accepted}} = \frac{c_{\mathrm{gen}}}{p}

Why this formula appears here

Rejection sampling generates multiple candidate completions per prompt from a model and keeps only the ones a separate scoring step accepts, using the survivors as supervised fine-tuning data. Llama 2’s RLHF pipeline used exactly this for its first four rounds, generating candidates from the 70-billion-parameter model, scoring them against a trained reward model, and only introducing proximal policy optimisation as a second mechanism in later rounds once the returns from sampling alone began to taper [ 6 ] . A second, independent demonstration in mathematical reasoning makes the mechanism even more explicit: rejection sampling fine-tuning collects correct reasoning paths directly from a…

Read the full article-specific guide →

Read the representative guide

cacceptedc_{\mathrm{accepted}}

Symbol c_accepted

cac_accepted is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →
cgenc_{\mathrm{gen}}

Symbol c_gen

cgc_gen occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

caccepted=cgenpc_{\mathrm{accepted}} = \frac{c_{\mathrm{gen}}}{p}

Equation 8 · Data & Training

Comparing the Main Approaches to Training Data and Synthetic Data

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Rejection sampling generates multiple candidate completions per prompt from a model and keeps only the ones a separate scoring step accepts, using the survivors as supervised fine-tuning data. Llama 2’s RLHF pipeline used exactly this for its first four rounds, generating candidates from the 70-billion-parameter model, scoring them against a trained reward model, and only introducing proximal policy optimisation as a second mechanism in later rounds once the returns from sampling alone began to taper [ 6 ] . A second, independent demonstration in mathematical reasoning makes the mechanism even more explicit: rejection sampling fine-tuning collects correct reasoning paths directly from a…

Meanings in this article

  • pp: the fraction.
Equation guide → · Article →