← Mathematical compendium

Published equation contexts

Attention(Q,K,V)=softmax ⁣(QK⊤dk)V\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V

Why this formula appears here

Vaswani and colleagues removed recurrence entirely, proposing an architecture “based solely on attention mechanisms, dispensing with recurrence and convolutions entirely”, and reported 28.4 BLEU on WMT 2014 English-to-German and 41.8 on English-to-French with substantially less training time [ 11 ] . The core operation is a single scaled dot-product: Attention(Q,K,V)=softmax ⁣(QK⊤dk)V\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V. The analytical point is that this is primarily a hardware result dressed as an architectural one. Every position attends to every other position in one parallel matrix multiplication, which maps precisely onto accelerator hardware in a way that a sequential recurrence never can. The transformer is what made it economically…

Read the full article-specific guide →

Read the representative guide

K⊤K^{\top}

Symbol K^top

KtK^top occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Read this term in its guide →
dkd_k

Symbol d_k

dkd_k occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Attention(Q,K,V)=softmax ⁣(QK⊤dk)V.\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V.

Equation 6 · Foundation Models

From n-Grams to Reasoning Models: A Technical History of the Language Model

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Vaswani and colleagues removed recurrence entirely, proposing an architecture “based solely on attention mechanisms, dispensing with recurrence and convolutions entirely”, and reported 28.4 BLEU on WMT 2014 English-to-German and 41.8 on English-to-French with substantially less training time [ 11 ] . The core operation is a single scaled dot-product: Attention(Q,K,V)=softmax ⁣(QK⊤dk)V\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V. The analytical point is that this is primarily a hardware result dressed as an architectural one. Every position attends to every other position in one parallel matrix multiplication, which maps precisely onto accelerator hardware in a way that a sequential recurrence never can. The transformer is what made it economically…

Equation guide → · Article →