← Back to article

Equation 6 · From n-Grams to Reasoning Models: A Technical History of the Language Model

What does this equation mean?

Attention(Q,K,V)=softmax ⁣(QK⊤dk)V.\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V.

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Start withQK^top
Divide bysqrtd_k
This relates toAttention(Q, K, V)
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

QQ

Symbol Q

Q occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

KK

Symbol K

K occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

VV

Symbol V

V is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

K⊤K^{\top}

Symbol K^top

KtK^top occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

dkd_k

Symbol d_k

dkd_k occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
√

√

Take a square root.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →
QK⊤QK^{\top}

Numerator: QK^top

The complete quantity above the fraction bar.

Understand this part →

dk\sqrt{d_k}

Denominator: sqrtd_k

The complete quantity below the fraction bar; it must be nonzero for this division.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Vaswani and colleagues removed recurrence entirely, proposing an architecture “based solely on attention mechanisms, dispensing with recurrence and convolutions entirely”, and reported 28.4 BLEU on WMT 2014 English-to-German and 41.8 on English-to-French with substantially less training time [ 11 ] . The core operation is a single scaled dot-product: Attention(Q,K,V)=softmax ⁣(QK⊤dk)V\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V. The analytical point is that this is primarily a hardware result dressed as an architectural one. Every position attends to every other position in one parallel matrix multiplication, which maps precisely onto accelerator hardware in a way that a sequential recurrence never can. The transformer is what made it economically…
Read the full surrounding passage
Vaswani and colleagues removed recurrence entirely, proposing an architecture “based solely on attention mechanisms, dispensing with recurrence and convolutions entirely”, and reported 28.4 BLEU on WMT 2014 English-to-German and 41.8 on English-to-French with substantially less training time [ 11 ] . The core operation is a single scaled dot-product: Attention(Q,K,V)=softmax ⁣(QK⊤dk)V\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V. The analytical point is that this is primarily a hardware result dressed as an architectural one. Every position attends to every other position in one parallel matrix multiplication, which maps precisely onto accelerator hardware in a way that a sequential recurrence never can. The transformer is what made it economically possible to spend the training budgets that the scaling era then spent. Interpretation, clearly labelled as such: without this step the scaling laws would still have been true and would still have been unaffordable.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to From n-Grams to Reasoning Models: A Technical History of the Language Model

See this formula across 1 published context →

Browse the mathematical compendium →