← All parts of this equation

Equation 1 · Part 14 · From n-Grams to Reasoning Models: A Technical History of the Language Model

Ending index or upper bound: T

P(w1,w2,…,wT)=∏t=1TP(wt∣w1,…,wt−1).P(w_1, w_2, \ldots, w_T) = \prod_{t=1}^{T} P(w_t \mid w_1, \ldots, w_{t-1}).
TT

What this part means

This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.

Its job in the formula

T appears in the bound of this product. The bound states where the repeated operation starts, ends, or which values it includes.

The passage around this formula

First, the object of study became prediction . A model of language is a probability assignment over what comes next. The chain rule makes this exact for any sequence of tokens: P(w1,w2,…,wT)=∏t=1TP(wt∣w1,…,wt−1)P(w_1, w_2, \ldots, w_T) = \prod_{t=1}^{T} P(w_t \mid w_1, \ldots, w_{t-1}). Second, quality became measurable without a task . Cross-entropy on held-out text is a number, and a lower number is unambiguously better. That gave the field a scalar to descend for the next seventy years, long before anyone knew what descending it would buy. Brown and colleagues later made the benchmark concrete, estimating an upper bound of 1.75 bits per character for English from a word trigram model measured against a balanced sample, and proposing a common corpus as a standard against which…

Read this part in the article →

Learn the underlying idea

Σ adds a collection of terms. Π multiplies them. The lower and upper labels tell you which terms belong to the collection.

Open the illustrated sums and products: repeat an operation over an index guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.