← All parts of this equation

Equation 22 · Part 2 · Locality Is the Whole Game: The Memory Hierarchy and What a Kernel Does Not Read

superscript

S=QK⊤\mathbf{S} = \mathbf{Q}\mathbf{K}^{\top}
superscript

What this part means

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Its job in the formula

A raised mark can be a power or an index. Its position and the surrounding notation determine which.

The passage around this formula

Standard attention forms the score matrix S\mathbf{S} = Q\mathbf{Q}K⊤\mathbf{K}^{\top} , applies a row-wise softmax to obtain P\mathbf{P} , and multiplies by V\mathbf{V} . Both S\mathbf{S} and P\mathbf{P} are N ×\times N , and standard implementations materialise them in main memory, which is quadratic in sequence length [ 1 ] . That materialisation is the problem. It is a large intermediate, it is written and read at least once, and the operations applied to it — masking, softmax, dropout — are exactly the memory-bound kind.

Read this part in the article →

Learn the underlying idea

An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.

Open the illustrated exponents: repeated multiplication and powers guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.