← All parts of this equation

Equation 8 · Part 3 · OpenAI Model Systems from First Principles: Weights, Post-Training, and Inference Compute

Symbol D

Cpre≈6ND,C_{\mathrm{pre}} \approx 6 N D,
DD

What this part means

D is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Its job in the formula

D is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

The passage around this formula

Pretraining compute for a dense transformer is well approximated by Cpre≈6NDC_{\mathrm{pre}} \approx 6 N D. with N parameters and D training tokens, the factor of six counting the forward and backward passes. Empirically, test loss falls as a power law in each of parameters, data, and compute over many orders of magnitude, a relationship first characterised systematically by Kaplan and colleagues [ 5 ] . The important structural feature is the functional form: a term of the shape

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.