← Mathematical compendium
Published equation contexts
S
Why this formula appears here
Standard attention forms the score matrix S = QK⊤ , applies a row-wise softmax to obtain P , and multiplies by V . Both S and P are N × N , and standard implementations materialise them in main memory, which is quadratic in sequence length [ 1 ] . That materialisation is the problem. It is a large intermediate, it is written and read at least once, and the operations applied to it — masking, softmax, dropout — are exactly the memory-bound kind.
Read the full article-specific guide →
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
Research cited beside this formula
Published contexts (1)
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 25 · AI Hardware & Semiconductors
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
Standard attention forms the score matrix S = QK⊤ , applies a row-wise softmax to obtain P , and multiplies by V . Both S and P are N × N , and standard implementations materialise them in main memory, which is quadratic in sequence length [ 1 ] . That materialisation is the problem. It is a large intermediate, it is written and read at least once, and the operations applied to it — masking, softmax, dropout — are exactly the memory-bound kind.
Equation guide → ·
Article →