← All parts of this equation

Equation 7 · Part 3 · From n-Grams to Reasoning Models: A Technical History of the Language Model

Symbol N_c

L(N)≈(NcN)αN.L(N) \approx \left(\frac{N_c}{N}\right)^{\alpha_N}.
NcN_c

What this part means

NcN_c occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Its job in the formula

NcN_c occurs above the fraction bar. The numerator is divided by the entire denominator below it.

The passage around this formula

By this point compute had become the binding constraint, and the field did not know how to spend it. Kaplan and colleagues characterised the relationship empirically, finding that cross-entropy loss scales as a power law in model size, dataset size and compute, with trends spanning more than seven orders of magnitude, and concluding that compute-efficient training meant very large models on relatively modest data, stopped well short of convergence [ 16 ] . The functional form is the important part: L(N)≈(NcN)αNL(N) \approx \left(\frac{N_c}{N}\right)^{\alpha_N}. Hoffmann and colleagues then trained over four hundred models and reached a different allocation: model size and training tokens should scale in roughly equal proportion, and…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.