← All parts of this equation

Equation 1 · Part 10 · From n-Grams to Reasoning Models: A Technical History of the Language Model

subtraction

P(w1,w2,…,wT)=∏t=1TP(wt∣w1,…,wt−1).P(w_1, w_2, \ldots, w_T) = \prod_{t=1}^{T} P(w_t \mid w_1, \ldots, w_{t-1}).
subtraction

What this part means

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Its job in the formula

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

The passage around this formula

First, the object of study became prediction . A model of language is a probability assignment over what comes next. The chain rule makes this exact for any sequence of tokens: P(w1,w2,…,wT)=∏t=1TP(wt∣w1,…,wt−1)P(w_1, w_2, \ldots, w_T) = \prod_{t=1}^{T} P(w_t \mid w_1, \ldots, w_{t-1}). Second, quality became measurable without a task . Cross-entropy on held-out text is a number, and a lower number is unambiguously better. That gave the field a scalar to descend for the next seventy years, long before anyone knew what descending it would buy. Brown and colleagues later made the benchmark concrete, estimating an upper bound of 1.75 bits per character for English from a word trigram model measured against a balanced sample, and proposing a common corpus as a standard against which…

Read this part in the article →

Learn the underlying idea

Addition combines quantities; subtraction measures the signed difference between them. Parentheses show what is combined before the rest of the expression is evaluated.

Open the illustrated addition and subtraction in an equation guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.