← Mathematical compendium

Published equation contexts

Cpre≈6NDC_{\mathrm{pre}} \approx 6ND

Why this formula appears here

The architecture underneath a Claude model, like nearly every other frontier language system, is a decoder-only transformer trained to predict the next token over a very large corpus [ 12 ] . What that training actually produces is not knowledge in any curated sense but a prior — a broad distribution over plausible continuations that encodes an enormous amount about language, code, and the structure of arguments, and comparatively little about what any particular operator wants done with it. Two variables set the price of that prior. Pretraining compute for a dense transformer is well approximated by Cpre≈6NDC_{\mathrm{pre}} \approx 6ND. where N is the parameter count and D is the number of training tokens,…

Read the full article-specific guide →

Read the representative guide

CpreC_{\mathrm{pre}}

Symbol C_pre

CpC_pre is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Its accuracy depends on the assumptions and range of use described in the article.

Research cited beside this formula

Published contexts (2)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Cpre≈6ND,C_{\mathrm{pre}} \approx 6ND,

Equation 1 · Foundation Models

Claude, From First Principles: Training, Constitutional Methods, and What Actually Shapes a Response

This equation gives an approximation: it relates the quantities while allowing an approximation.

The architecture underneath a Claude model, like nearly every other frontier language system, is a decoder-only transformer trained to predict the next token over a very large corpus [ 12 ] . What that training actually produces is not knowledge in any curated sense but a prior — a broad distribution over plausible continuations that encodes an enormous amount about language, code, and the structure of arguments, and comparatively little about what any particular operator wants done with it. Two variables set the price of that prior. Pretraining compute for a dense transformer is well approximated by Cpre≈6NDC_{\mathrm{pre}} \approx 6ND. where N is the parameter count and D is the number of training tokens,…

Meanings in this article

  • NN: the parameter count.
  • DD: the number of training tokens, the factor of six accounting for the forward and backward pass together.
Equation guide → · Article →
Cpre≈6ND,C_{\mathrm{pre}} \approx 6 N D,

Equation 8 · Foundation Models

OpenAI Model Systems from First Principles: Weights, Post-Training, and Inference Compute

This equation gives an approximation: it relates the quantities while allowing an approximation.

Pretraining compute for a dense transformer is well approximated by Cpre≈6NDC_{\mathrm{pre}} \approx 6 N D. with N parameters and D training tokens, the factor of six counting the forward and backward passes. Empirically, test loss falls as a power law in each of parameters, data, and compute over many orders of magnitude, a relationship first characterised systematically by Kaplan and colleagues [ 5 ] . The important structural feature is the functional form: a term of the shape

Equation guide → · Article →