← All parts of this equation

Equation 11 · Part 10 · OpenAI Model Systems from First Principles: Weights, Post-Training, and Inference Compute

superscript

L(N)≈L∞+(NcN)αNL(N) \approx L_\infty + \left(\frac{N_c}{N}\right)^{\alpha_N}
superscript

What this part means

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Its job in the formula

A raised mark can be a power or an index. Its position and the surrounding notation determine which.

The passage around this formula

with N parameters and D training tokens, the factor of six counting the forward and backward passes. Empirically, test loss falls as a power law in each of parameters, data, and compute over many orders of magnitude, a relationship first characterised systematically by Kaplan and colleagues [ 5 ] . The important structural feature is the functional form: a term of the shape L(N)≈L∞+(NcN)αNL(N) \approx L_\infty + \left(\frac{N_c}{N}\right)^{\alpha_N}. has an irreducible floor L∞L_\infty and diminishing returns above it. Doubling N buys a fixed decrement in loss, not a fixed multiple of capability.

Read this part in the article →

Learn the underlying idea

An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.

Open the illustrated exponents: repeated multiplication and powers guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.