← All parts of this equation

Equation 2 · Part 6 · From n-Grams to Reasoning Models: A Technical History of the Language Model

≈

P(wt∣w1,…,wt−1)≈P(wt∣wt−n+1,…,wt−1).P(w_t \mid w_1, \ldots, w_{t-1}) \approx P(w_t \mid w_{t-n+1}, \ldots, w_{t-1}).
≈

What this part means

Approximately equal to; the equality is not exact.

Its job in the formula

Approximately equal to; the equality is not exact.

The passage around this formula

The obvious implementation of the chain rule is to condition on everything. That is impossible, so the Markov approximation truncates the history to a window of fixed width: P(wt∣w1,…,wt−1)≈P(wt∣wt−n+1,…,wt−1)P(w_t \mid w_1, \ldots, w_{t-1}) \approx P(w_t \mid w_{t-n+1}, \ldots, w_{t-1}). Estimate each conditional by counting. This is the n-gram model, and it dominated applied language modelling for roughly three decades because it is cheap, transparent, and surprisingly hard to beat on enough data.

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.