← Back to article

Equation 10 · OpenAI Model Systems from First Principles: Weights, Post-Training, and Inference Compute

What does this equation mean?

DD

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

DD

Symbol D

D is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

with N parameters and D training tokens, the factor of six counting the forward and backward passes. Empirically, test loss falls as a power law in each of parameters, data, and compute over many orders of magnitude, a relationship first characterised systematically by Kaplan and colleagues [ 5 ] . The important structural feature is the functional form: a term of the shape

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to OpenAI Model Systems from First Principles: Weights, Post-Training, and Inference Compute

Browse the mathematical compendium →