← Mathematical compendium

Published equation contexts

Ctotal=Cpre+Cpost+Q⋅cˉinfC_{\mathrm{total}} = C_{\mathrm{pre}} + C_{\mathrm{post}} + Q \cdot \bar{c}_{\mathrm{inf}}

Add the one-time compute used to train and refine the model to the compute spent answering requests over its lifetime.

Why this formula appears here

Write the total lifetime computation of a deployed system as Ctotal=Cpre+Cpost+Q⋅cˉinfC_{\mathrm{total}} = C_{\mathrm{pre}} + C_{\mathrm{post}} + Q \cdot \bar{c}_{\mathrm{inf}}. where CpreC_{\mathrm{pre}} is pretraining compute, CpostC_{\mathrm{post}} is post-training compute, Q is the number of served requests over the system’s life, and cˉinf\bar{c}_{\mathrm{inf}} is the mean compute per request. Until roughly 2024 the third term was treated as approximately fixed for a given model, and public discussion of capability collapsed onto the first. That assumption no longer holds. On current OpenAI models, cˉinf\bar{c}_{\mathrm{inf}} is a parameter the caller sets: as verified on 8 August 2026, the model guidance documents a reasoningeg_effort control taking the values none , low , medium , high , xhigh , and…

Read the full article-specific guide →

Read the representative guide

CpostC_{\mathrm{post}}

Post-training compute

The one-time computation used after pretraining to shape the model’s behavior.

Read this term in its guide →
cˉinf\bar{c}_{\mathrm{inf}}

Compute per request

The average computation used to answer one request. The bar means average.

Read this term in its guide →
Q⋅cˉinfQ \cdot \bar{c}_{\mathrm{inf}}

Lifetime inference compute

Requests multiplied by average compute per request.

Read this term in its guide →

How to interpret it

The last term grows with usage. Doubling the number of requests doubles that term if average compute per request stays the same. The equation is an accounting model, not a claim that every request costs the same.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Ctotal=Cpre+Cpost+Q⋅cˉinf,C_{\mathrm{total}} = C_{\mathrm{pre}} + C_{\mathrm{post}} + Q \cdot \bar{c}_{\mathrm{inf}},

Equation 1 · Foundation Models

OpenAI Model Systems from First Principles: Weights, Post-Training, and Inference Compute

The total compute bill has three parts: pretraining, post-training, and all the requests served afterwards.

Write the total lifetime computation of a deployed system as Ctotal=Cpre+Cpost+Q⋅cˉinfC_{\mathrm{total}} = C_{\mathrm{pre}} + C_{\mathrm{post}} + Q \cdot \bar{c}_{\mathrm{inf}}. where CpreC_{\mathrm{pre}} is pretraining compute, CpostC_{\mathrm{post}} is post-training compute, Q is the number of served requests over the system’s life, and cˉinf\bar{c}_{\mathrm{inf}} is the mean compute per request. Until roughly 2024 the third term was treated as approximately fixed for a given model, and public discussion of capability collapsed onto the first. That assumption no longer holds. On current OpenAI models, cˉinf\bar{c}_{\mathrm{inf}} is a parameter the caller sets: as verified on 8 August 2026, the model guidance documents a reasoningeg_effort control taking the values none , low , medium , high , xhigh , and…

Meanings in this article

Equation guide → · Article →