Equation 27 · OpenAI Model Systems from First Principles: Weights, Post-Training, and Inference Compute
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol k
k is part of the quantity the equation computes from the expression on the right.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The parallel-sampling case makes the shape of the returns explicit. If a single attempt succeeds with probability p and attempts were independent, the probability that at least one of k succeeds is . which is concave in k and saturates quickly. Two caveats destroy any naive extrapolation from it. Attempts from one model on one prompt are strongly correlated, so realised gains fall well below this bound; and is only achievable if something can identify the successful attempt. Without a verifier, extra samples buy candidates, not answers. This is precisely why the reasoning-effort control and the availability of parallel test-time compute are architectural facts about a…
Read the full surrounding passage
The parallel-sampling case makes the shape of the returns explicit. If a single attempt succeeds with probability p and attempts were independent, the probability that at least one of k succeeds is . which is concave in k and saturates quickly. Two caveats destroy any naive extrapolation from it. Attempts from one model on one prompt are strongly correlated, so realised gains fall well below this bound; and is only achievable if something can identify the successful attempt. Without a verifier, extra samples buy candidates, not answers. This is precisely why the reasoning-effort control and the availability of parallel test-time compute are architectural facts about a product rather than mere quality dials — the GPT-5 system card describes a variant that makes use of parallel test-time compute, and states that all models were evaluated at high reasoning effort [ 1 ] .
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to OpenAI Model Systems from First Principles: Weights, Post-Training, and Inference Compute