Total compute
All computation spent on this deployed model system over its lifetime.
Read this term in its guide →Published equation contexts
Add the one-time compute used to train and refine the model to the compute spent answering requests over its lifetime.
Write the total lifetime computation of a deployed system as . where is pretraining compute, is post-training compute, Q is the number of served requests over the system’s life, and is the mean compute per request. Until roughly 2024 the third term was treated as approximately fixed for a given model, and public discussion of capability collapsed onto the first. That assumption no longer holds. On current OpenAI models, is a parameter the caller sets: as verified on 8 August 2026, the model guidance documents a reasoninffort control taking the values none , low , medium , high , xhigh , and…
All computation spent on this deployed model system over its lifetime.
Read this term in its guide →The one-time computation used to learn from the initial training data.
Read this term in its guide →The one-time computation used after pretraining to shape the model’s behavior.
Read this term in its guide →How many requests the system serves during the period being counted.
Read this term in its guide →The average computation used to answer one request. The bar means average.
Read this term in its guide →Requests multiplied by average compute per request.
Read this term in its guide →The last term grows with usage. Doubling the number of requests doubles that term if average compute per request stays the same. The equation is an accounting model, not a claim that every request costs the same.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 1 · Foundation Models
The total compute bill has three parts: pretraining, post-training, and all the requests served afterwards.
Write the total lifetime computation of a deployed system as . where is pretraining compute, is post-training compute, Q is the number of served requests over the system’s life, and is the mean compute per request. Until roughly 2024 the third term was treated as approximately fixed for a given model, and public discussion of capability collapsed onto the first. That assumption no longer holds. On current OpenAI models, is a parameter the caller sets: as verified on 8 August 2026, the model guidance documents a reasoninffort control taking the values none , low , medium , high , xhigh , and…