Equation 22 · Locality Is the Whole Game: The Memory Hierarchy and What a Kernel Does Not Read
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Standard attention forms the score matrix = , applies a row-wise softmax to obtain , and multiplies by . Both and are N N , and standard implementations materialise them in main memory, which is quadratic in sequence length [ 1 ] . That materialisation is the problem. It is a large intermediate, it is written and read at least once, and the operations applied to it — masking, softmax, dropout — are exactly the memory-bound kind.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to Locality Is the Whole Game: The Memory Hierarchy and What a Kernel Does Not Read