← All parts of this equation

Equation 6 · Part 4 · From n-Grams to Reasoning Models: A Technical History of the Language Model

Symbol K^top

Attention(Q,K,V)=softmax ⁣(QK⊤dk)V.\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V.
K⊤K^{\top}

What this part means

KtK^top occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Its job in the formula

KtK^top occurs above the fraction bar. The numerator is divided by the entire denominator below it.

The passage around this formula

Vaswani and colleagues removed recurrence entirely, proposing an architecture “based solely on attention mechanisms, dispensing with recurrence and convolutions entirely”, and reported 28.4 BLEU on WMT 2014 English-to-German and 41.8 on English-to-French with substantially less training time [ 11 ] . The core operation is a single scaled dot-product: Attention(Q,K,V)=softmax ⁣(QK⊤dk)V\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V. The analytical point is that this is primarily a hardware result dressed as an architectural one. Every position attends to every other position in one parallel matrix multiplication, which maps precisely onto accelerator hardware in a way that a sequential recurrence never can. The transformer is what made it economically…

Read this part in the article →

Learn the underlying idea

An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.

Open the illustrated exponents: repeated multiplication and powers guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.