← All parts of this equation

Equation 6 · Part 5 · From n-Grams to Reasoning Models: A Technical History of the Language Model

Symbol d_k

Attention(Q,K,V)=softmax ⁣(QK⊤dk)V.\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V.
dkd_k

What this part means

dkd_k occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Its job in the formula

dkd_k occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

The passage around this formula

Vaswani and colleagues removed recurrence entirely, proposing an architecture “based solely on attention mechanisms, dispensing with recurrence and convolutions entirely”, and reported 28.4 BLEU on WMT 2014 English-to-German and 41.8 on English-to-French with substantially less training time [ 11 ] . The core operation is a single scaled dot-product: Attention(Q,K,V)=softmax ⁣(QK⊤dk)V\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V. The analytical point is that this is primarily a hardware result dressed as an architectural one. Every position attends to every other position in one parallel matrix multiplication, which maps precisely onto accelerator hardware in a way that a sequential recurrence never can. The transformer is what made it economically…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.