← All parts of this equation

Equation 2 · Part 9 · From n-Grams to Reasoning Models: A Technical History of the Language Model

subscript

P(wt∣w1,…,wt−1)≈P(wt∣wt−n+1,…,wt−1).P(w_t \mid w_1, \ldots, w_{t-1}) \approx P(w_t \mid w_{t-n+1}, \ldots, w_{t-1}).
subscript

What this part means

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Its job in the formula

A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.

The passage around this formula

The obvious implementation of the chain rule is to condition on everything. That is impossible, so the Markov approximation truncates the history to a window of fixed width: P(wt∣w1,…,wt−1)≈P(wt∣wt−n+1,…,wt−1)P(w_t \mid w_1, \ldots, w_{t-1}) \approx P(w_t \mid w_{t-n+1}, \ldots, w_{t-1}). Estimate each conditional by counting. This is the n-gram model, and it dominated applied language modelling for roughly three decades because it is cheap, transparent, and surprisingly hard to beat on enough data.

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.