← Mathematical compendium
Published equation contexts
S=QK⊤
Why this formula appears here
Standard attention forms the score matrix S = QK⊤ , applies a row-wise softmax to obtain P , and multiplies by V . Both S and P are N × N , and standard implementations materialise them in main memory, which is quadratic in sequence length [ 1 ] . That materialisation is the problem. It is a large intermediate, it is written and read at least once, and the operations applied to it — masking, softmax, dropout — are exactly the memory-bound kind.
Read the full article-specific guide →
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
Research cited beside this formula
Published contexts (1)
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 22 · AI Hardware & Semiconductors
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
Standard attention forms the score matrix S = QK⊤ , applies a row-wise softmax to obtain P , and multiplies by V . Both S and P are N × N , and standard implementations materialise them in main memory, which is quadratic in sequence length [ 1 ] . That materialisation is the problem. It is a large intermediate, it is written and read at least once, and the operations applied to it — masking, softmax, dropout — are exactly the memory-bound kind.
Equation guide → ·
Article →