← All parts of this equation

Equation 6 · Part 2 · From n-Grams to Reasoning Models: A Technical History of the Language Model

Symbol K

Attention(Q,K,V)=softmax ⁣(QK⊤dk)V.\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V.
KK

What this part means

K occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Its job in the formula

K occurs above the fraction bar. The numerator is divided by the entire denominator below it.

The passage around this formula

Vaswani and colleagues removed recurrence entirely, proposing an architecture “based solely on attention mechanisms, dispensing with recurrence and convolutions entirely”, and reported 28.4 BLEU on WMT 2014 English-to-German and 41.8 on English-to-French with substantially less training time [ 11 ] . The core operation is a single scaled dot-product: Attention(Q,K,V)=softmax ⁣(QK⊤dk)V\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V. The analytical point is that this is primarily a hardware result dressed as an architectural one. Every position attends to every other position in one parallel matrix multiplication, which maps precisely onto accelerator hardware in a way that a sequential recurrence never can. The transformer is what made it economically…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.