← Mathematical compendium

Published equation contexts

q(b)q(b)

Why this formula appears here

Consider a task instance with a policy πθ\pi_\theta and a compute budget b spent at inference, whether as sequential deliberation, parallel sampling, or search against a verifier. Expected quality is some q(b) that rises and saturates. Snell and colleagues studied this directly and reported that allocating test-time compute adaptively to the difficulty of the prompt substantially outperforms uniform allocation, and that in some regimes additional inference compute is a more effective use of a marginal FLOP than additional parameters [ 9 ] . The practically important half of that finding is the first: the optimal b is a function of the instance, not of the model.

Read the full article-specific guide →

Read the representative guide

qq

Symbol q

q is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
bb

Symbol b

b is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

q(b)q(b)

Equation 23 · Foundation Models

OpenAI Model Systems from First Principles: Weights, Post-Training, and Inference Compute

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

Consider a task instance with a policy πθ\pi_\theta and a compute budget b spent at inference, whether as sequential deliberation, parallel sampling, or search against a verifier. Expected quality is some q(b) that rises and saturates. Snell and colleagues studied this directly and reported that allocating test-time compute adaptively to the difficulty of the prompt substantially outperforms uniform allocation, and that in some regimes additional inference compute is a more effective use of a marginal FLOP than additional parameters [ 9 ] . The practically important half of that finding is the first: the optimal b is a function of the instance, not of the model.

Equation guide → · Article →