← Mathematical compendium

Published equation contexts

P(wt∣w1,…,wt−1)≈P(wt∣wt−n+1,…,wt−1)P(w_t \mid w_1, \ldots, w_{t-1}) \approx P(w_t \mid w_{t-n+1}, \ldots, w_{t-1})

Why this formula appears here

The obvious implementation of the chain rule is to condition on everything. That is impossible, so the Markov approximation truncates the history to a window of fixed width: P(wt∣w1,…,wt−1)≈P(wt∣wt−n+1,…,wt−1)P(w_t \mid w_1, \ldots, w_{t-1}) \approx P(w_t \mid w_{t-n+1}, \ldots, w_{t-1}). Estimate each conditional by counting. This is the n-gram model, and it dominated applied language modelling for roughly three decades because it is cheap, transparent, and surprisingly hard to beat on enough data.

Read the full article-specific guide →

Read the representative guide

PP

Symbol P

P is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
wtw_t

Symbol w_t

wtw_t is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
w1w_1

Symbol w_1

w1w_1 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
wt−1w_{t-1}

Symbol w_t-1

wtw_t-1 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
wt−n+1w_{t-n+1}

Symbol w_t-n+1

wtw_t-n+1 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Its accuracy depends on the assumptions and range of use described in the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

P(wt∣w1,…,wt−1)≈P(wt∣wt−n+1,…,wt−1).P(w_t \mid w_1, \ldots, w_{t-1}) \approx P(w_t \mid w_{t-n+1}, \ldots, w_{t-1}).

Equation 2 · Foundation Models

From n-Grams to Reasoning Models: A Technical History of the Language Model

This equation gives an approximation: it relates the quantities while allowing an approximation.

The obvious implementation of the chain rule is to condition on everything. That is impossible, so the Markov approximation truncates the history to a window of fixed width: P(wt∣w1,…,wt−1)≈P(wt∣wt−n+1,…,wt−1)P(w_t \mid w_1, \ldots, w_{t-1}) \approx P(w_t \mid w_{t-n+1}, \ldots, w_{t-1}). Estimate each conditional by counting. This is the n-gram model, and it dominated applied language modelling for roughly three decades because it is cheap, transparent, and surprisingly hard to beat on enough data.

Equation guide → · Article →