← All parts of this equation

Equation 1 · Part 4 · From n-Grams to Reasoning Models: A Technical History of the Language Model

Symbol w_T

P(w1,w2,…,wT)=∏t=1TP(wt∣w1,…,wt−1).P(w_1, w_2, \ldots, w_T) = \prod_{t=1}^{T} P(w_t \mid w_1, \ldots, w_{t-1}).
wTw_T

What this part means

wTw_T is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Its job in the formula

wTw_T is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

The passage around this formula

First, the object of study became prediction . A model of language is a probability assignment over what comes next. The chain rule makes this exact for any sequence of tokens: P(w1,w2,…,wT)=∏t=1TP(wt∣w1,…,wt−1)P(w_1, w_2, \ldots, w_T) = \prod_{t=1}^{T} P(w_t \mid w_1, \ldots, w_{t-1}). Second, quality became measurable without a task . Cross-entropy on held-out text is a number, and a lower number is unambiguously better. That gave the field a scalar to descend for the next seventy years, long before anyone knew what descending it would buy. Brown and colleagues later made the benchmark concrete, estimating an upper bound of 1.75 bits per character for English from a word trigram model measured against a balanced sample, and proposing a common corpus as a standard against which…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.