← All parts of this equation

Equation 1 · Part 4 · How Training Data and Synthetic Data Actually Work

Symbol p_fail

Ccurriculum=N⋅(cgenerate+pfail⋅cregenerate)C_{\text{curriculum}} = N \cdot \left( c_{\text{generate}} + p_{\text{fail}} \cdot c_{\text{regenerate}} \right)
pfailp_{\text{fail}}

What this part means

the fraction that fail correctness or contamination filtering and must be regenerated.

Its job in the formula

pfp_fail is one of the signed contributions combined to compute the quantity on the left.

Where the article explains it

where N is the number of exercises targeted, cgeneratec_{\text{generate}} is the compute cost of producing one candidate item, pfailp_{\text{fail}} is the fraction that fail correctness or contamination filtering and must be regenerated, and cregeneratec_{\text{regenerate}} is the cost of another attempt.

The passage around this formula

…reason about scale is a simple accounting identity rather than a law of nature: Ccurriculum=N⋅(cgenerate+pfail⋅cregenerate)C_{\text{curriculum}} = N \cdot \left( c_{\text{generate}} + p_{\text{fail}} \cdot c_{\text{regenerate}} \right). where N is the number of exercises targeted, cgeneratec_{\text{generate}} is the compute cost of producing one candidate item, pfailp_{\text{fail}} is the fraction that fail correctness or contamination filtering and must be regenerated, and cregeneratec_{\text{regenerate}} is the cost of another attempt. This is not drawn from a specific paper’s disclosed cost model — none of the sources here publish exact figures — it is included only…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.