← Back to article

Equation 11 · OpenAI Model Systems from First Principles: Weights, Post-Training, and Inference Compute

What does this equation mean?

L(N)≈L∞+(NcN)αNL(N) \approx L_\infty + \left(\frac{N_c}{N}\right)^{\alpha_N}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

LL

Symbol L

L is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

NN

Symbol N

N is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

L∞L_\infty

Symbol L_infty

the irreducible floor.

Understand this part →

NcN_c

Symbol N_c

NcN_c occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

αN\alpha_N

Symbol alpha_N

alphaNa_N is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
≈

≈

Approximately equal to; the equality is not exact.

Understand this part →

addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.

What the article says around this equation

with N parameters and D training tokens, the factor of six counting the forward and backward passes. Empirically, test loss falls as a power law in each of parameters, data, and compute over many orders of magnitude, a relationship first characterised systematically by Kaplan and colleagues [ 5 ] . The important structural feature is the functional form: a term of the shape L(N)≈L∞+(NcN)αNL(N) \approx L_\infty + \left(\frac{N_c}{N}\right)^{\alpha_N}. has an irreducible floor L∞L_\infty and diminishing returns above it. Doubling N buys a fixed decrement in loss, not a fixed multiple of capability.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to OpenAI Model Systems from First Principles: Weights, Post-Training, and Inference Compute

See this formula across 1 published context →

Browse the mathematical compendium →