← Mathematical compendium

Published equation contexts

Pr⁡(correct∣c^=c)≈c\Pr(\text{correct} \mid \hat{c} = c) \approx c

Why this formula appears here

The structural difficulty is that a next-token distribution is not an epistemic state. A model trained to maximise likelihood over text produces a confident-sounding continuation because confident-sounding continuations are what the corpus contains, not because it has assessed its own evidence. Calibration can be measured — for a predicted confidence c one can ask whether Pr⁡(correct∣c^=c)≈c\Pr(\text{correct} \mid \hat{c} = c) \approx c. holds across bins — and it can be improved by post-hoc adjustment. What has not been demonstrated is a mechanism by which a model represents its own ignorance in a way that survives fine-tuning, distribution shift, and the pressure of an objective that rewards answering.

Read the full article-specific guide →

Read the representative guide

c^\hat{c}

Symbol hatc

hatc appears in the conditional probability being evaluated. The vertical bar identifies the information or condition supplied to that probability.

Read this term in its guide →
cc

Symbol c

c appears in the conditional probability being evaluated. The vertical bar identifies the information or condition supplied to that probability.

Read this term in its guide →
Pr⁡\Pr

Probability operator

The probability operator gives the chance of the event named inside its brackets or parentheses.

Read this term in its guide →

How to interpret it

Its accuracy depends on the assumptions and range of use described in the article. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Pr⁡(correct∣c^=c)≈c\Pr(\text{correct} \mid \hat{c} = c) \approx c

Equation 2 · Foundation Models

What We Still Cannot Do: Open Problems in Frontier Model Systems

This equation gives an approximation: it relates the quantities while allowing an approximation.

The structural difficulty is that a next-token distribution is not an epistemic state. A model trained to maximise likelihood over text produces a confident-sounding continuation because confident-sounding continuations are what the corpus contains, not because it has assessed its own evidence. Calibration can be measured — for a predicted confidence c one can ask whether Pr⁡(correct∣c^=c)≈c\Pr(\text{correct} \mid \hat{c} = c) \approx c. holds across bins — and it can be improved by post-hoc adjustment. What has not been demonstrated is a mechanism by which a model represents its own ignorance in a way that survives fine-tuning, distribution shift, and the pressure of an objective that rewards answering.

Equation guide → · Article →