← All parts of this equation

Equation 8 · Part 9 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

fraction

L(N,D)≈E+ANα+BDβL(N, D) \approx E + \frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}}
fraction

What this part means

Divide the expression above the line by the one below it.

Its job in the formula

The expression above the fraction bar is divided by the complete expression below it. The denominator must not be zero.

The passage around this formula

Its intellectual foundation is the same scaling-law literature that shaped how large models are trained, applied in the opposite direction. Hoffmann and colleagues showed that contemporary large language models had been trained on too little data relative to their parameter count, and that for a fixed training budget, model size and training tokens should grow in roughly equal proportion — “for every doubling of model size the number of training tokens should also be doubled” [ 4 ] . The underlying loss law is commonly written as L(N,D)≈E+ANα+BDβL(N, D) \approx E + \frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}}. with N parameters and D training tokens. Chinchilla’s question was: given a fixed compute budget C , what N and D minimize L ? The…

Read this part in the article →

Learn the underlying idea

A fraction a/b means a divided by b. The top number is the numerator; the bottom number is the denominator, and it cannot be zero.

Open the illustrated fractions: division written vertically guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.