Symbol C_pre
re is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Published equation contexts
The architecture underneath a Claude model, like nearly every other frontier language system, is a decoder-only transformer trained to predict the next token over a very large corpus [ 12 ] . What that training actually produces is not knowledge in any curated sense but a prior — a broad distribution over plausible continuations that encodes an enormous amount about language, code, and the structure of arguments, and comparatively little about what any particular operator wants done with it. Two variables set the price of that prior. Pretraining compute for a dense transformer is well approximated by . where N is the parameter count and D is the number of training tokens,…
re is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →the number of training tokens, the factor of six accounting for the forward and backward pass together.
Read this term in its guide →Its accuracy depends on the assumptions and range of use described in the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 1 · Foundation Models
This equation gives an approximation: it relates the quantities while allowing an approximation.
The architecture underneath a Claude model, like nearly every other frontier language system, is a decoder-only transformer trained to predict the next token over a very large corpus [ 12 ] . What that training actually produces is not knowledge in any curated sense but a prior — a broad distribution over plausible continuations that encodes an enormous amount about language, code, and the structure of arguments, and comparatively little about what any particular operator wants done with it. Two variables set the price of that prior. Pretraining compute for a dense transformer is well approximated by . where N is the parameter count and D is the number of training tokens,…
Equation 8 · Foundation Models
This equation gives an approximation: it relates the quantities while allowing an approximation.
Pretraining compute for a dense transformer is well approximated by . with N parameters and D training tokens, the factor of six counting the forward and backward passes. Empirically, test loss falls as a power law in each of parameters, data, and compute over many orders of magnitude, a relationship first characterised systematically by Kaplan and colleagues [ 5 ] . The important structural feature is the functional form: a term of the shape
Equation guide → · Article →