← All parts of this equation

Equation 2 · Part 7 · From n-Grams to Reasoning Models: A Technical History of the Language Model

addition

P(wt∣w1,…,wt−1)≈P(wt∣wt−n+1,…,wt−1).P(w_t \mid w_1, \ldots, w_{t-1}) \approx P(w_t \mid w_{t-n+1}, \ldots, w_{t-1}).
addition

What this part means

Add the term after the plus sign to the term or group before it.

Its job in the formula

Add the term after the plus sign to the term or group before it.

The passage around this formula

The obvious implementation of the chain rule is to condition on everything. That is impossible, so the Markov approximation truncates the history to a window of fixed width: P(wt∣w1,…,wt−1)≈P(wt∣wt−n+1,…,wt−1)P(w_t \mid w_1, \ldots, w_{t-1}) \approx P(w_t \mid w_{t-n+1}, \ldots, w_{t-1}). Estimate each conditional by counting. This is the n-gram model, and it dominated applied language modelling for roughly three decades because it is cheap, transparent, and surprisingly hard to beat on enough data.

Read this part in the article →

Learn the underlying idea

Addition combines quantities; subtraction measures the signed difference between them. Parentheses show what is combined before the rest of the expression is evaluated.

Open the illustrated addition and subtraction in an equation guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.