← Mathematical compendium

Published equation contexts

V\mathbf{V}

Why this formula appears here

Standard attention forms the score matrix S\mathbf{S} = Q\mathbf{Q}K⊤\mathbf{K}^{\top} , applies a row-wise softmax to obtain P\mathbf{P} , and multiplies by V\mathbf{V} . Both S\mathbf{S} and P\mathbf{P} are N ×\times N , and standard implementations materialise them in main memory, which is quadratic in sequence length [ 1 ] . That materialisation is the problem. It is a large intermediate, it is written and read at least once, and the operations applied to it — masking, softmax, dropout — are exactly the memory-bound kind.

Read the full article-specific guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

V\mathbf{V}

Equation 24 · AI Hardware & Semiconductors

Locality Is the Whole Game: The Memory Hierarchy and What a Kernel Does Not Read

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

Standard attention forms the score matrix S\mathbf{S} = Q\mathbf{Q}K⊤\mathbf{K}^{\top} , applies a row-wise softmax to obtain P\mathbf{P} , and multiplies by V\mathbf{V} . Both S\mathbf{S} and P\mathbf{P} are N ×\times N , and standard implementations materialise them in main memory, which is quadratic in sequence length [ 1 ] . That materialisation is the problem. It is a large intermediate, it is written and read at least once, and the operations applied to it — masking, softmax, dropout — are exactly the memory-bound kind.

Equation guide → · Article →