← Mathematical compendium

Published equation contexts

L(N)≈L∞+(NcN)αNL(N) \approx L_\infty + \left(\frac{N_c}{N}\right)^{\alpha_N}

Why this formula appears here

with N parameters and D training tokens, the factor of six counting the forward and backward passes. Empirically, test loss falls as a power law in each of parameters, data, and compute over many orders of magnitude, a relationship first characterised systematically by Kaplan and colleagues [ 5 ] . The important structural feature is the functional form: a term of the shape L(N)≈L∞+(NcN)αNL(N) \approx L_\infty + \left(\frac{N_c}{N}\right)^{\alpha_N}. has an irreducible floor L∞L_\infty and diminishing returns above it. Doubling N buys a fixed decrement in loss, not a fixed multiple of capability.

Read the full article-specific guide →

Read the representative guide

LL

Symbol L

L is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
NN

Symbol N

N is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
αN\alpha_N

Symbol alpha_N

alphaNa_N is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

L(N)≈L∞+(NcN)αNL(N) \approx L_\infty + \left(\frac{N_c}{N}\right)^{\alpha_N}

Equation 11 · Foundation Models

OpenAI Model Systems from First Principles: Weights, Post-Training, and Inference Compute

This equation gives an approximation: it relates the quantities while allowing an approximation.

with N parameters and D training tokens, the factor of six counting the forward and backward passes. Empirically, test loss falls as a power law in each of parameters, data, and compute over many orders of magnitude, a relationship first characterised systematically by Kaplan and colleagues [ 5 ] . The important structural feature is the functional form: a term of the shape L(N)≈L∞+(NcN)αNL(N) \approx L_\infty + \left(\frac{N_c}{N}\right)^{\alpha_N}. has an irreducible floor L∞L_\infty and diminishing returns above it. Doubling N buys a fixed decrement in loss, not a fixed multiple of capability.

Meanings in this article

Equation guide → · Article →