← Back to article

Equation 1 · From n-Grams to Reasoning Models: A Technical History of the Language Model

What does this equation mean?

P(w1,w2,…,wT)=∏t=1TP(wt∣w1,…,wt−1).P(w_1, w_2, \ldots, w_T) = \prod_{t=1}^{T} P(w_t \mid w_1, \ldots, w_{t-1}).

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationsprod_t=1^T P(w_t mid w_1, ldots, w_t-1)
Result or conditionP(w_1, w_2, ldots, w_T)
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

PP

Symbol P

P is part of the quantity the equation computes from the expression on the right.

Understand this part →

w1w_1

Symbol w_1

w1w_1 is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

w2w_2

Symbol w_2

w2w_2 is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

wTw_T

Symbol w_T

wTw_T is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

tt

Symbol t

t appears in the bound of this product. The bound states where the repeated operation starts, ends, or which values it includes.

Understand this part →

TT

Symbol T

T appears in the bound of this product. The bound states where the repeated operation starts, ends, or which values it includes.

Understand this part →

wtw_t

Symbol w_t

wtw_t is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

wt−1w_{t-1}

Symbol w_t-1

wtw_t-1 is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
subtraction

subtraction

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →
t=1t=1

Starting index or lower bound: t=1

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Understand this part →

TT

Ending index or upper bound: T

This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

First, the object of study became prediction . A model of language is a probability assignment over what comes next. The chain rule makes this exact for any sequence of tokens: P(w1,w2,…,wT)=∏t=1TP(wt∣w1,…,wt−1)P(w_1, w_2, \ldots, w_T) = \prod_{t=1}^{T} P(w_t \mid w_1, \ldots, w_{t-1}). Second, quality became measurable without a task . Cross-entropy on held-out text is a number, and a lower number is unambiguously better. That gave the field a scalar to descend for the next seventy years, long before anyone knew what descending it would buy. Brown and colleagues later made the benchmark concrete, estimating an upper bound of 1.75 bits per character for English from a word trigram model measured against a balanced sample, and proposing a common corpus as a standard against which…
Read the full surrounding passage
First, the object of study became prediction . A model of language is a probability assignment over what comes next. The chain rule makes this exact for any sequence of tokens: P(w1,w2,…,wT)=∏t=1TP(wt∣w1,…,wt−1)P(w_1, w_2, \ldots, w_T) = \prod_{t=1}^{T} P(w_t \mid w_1, \ldots, w_{t-1}). Second, quality became measurable without a task . Cross-entropy on held-out text is a number, and a lower number is unambiguously better. That gave the field a scalar to descend for the next seventy years, long before anyone knew what descending it would buy. Brown and colleagues later made the benchmark concrete, estimating an upper bound of 1.75 bits per character for English from a word trigram model measured against a balanced sample, and proposing a common corpus as a standard against which to measure progress [ 3 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to From n-Grams to Reasoning Models: A Technical History of the Language Model

See this formula across 1 published context →

Browse the mathematical compendium →