Equation 28 · Locality Is the Whole Game: The Memory Hierarchy and What a Kernel Does Not Read
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol O
O is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol M^-1
1 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
FlashAttention converted that observation into wall-clock speed by combining tiling with recomputation. Blocks of the inputs are loaded into on-chip SRAM, the softmax reduction is performed incrementally across blocks using the online normaliser, and the output is written back without the score matrix ever reaching main memory; for the backward pass, the softmax normalisation factors are kept from the forward pass so that attention can be recomputed on-chip, which the authors state is faster than reading the stored intermediate back from HBM. Their analysis gives O( ) HBM accesses against for standard attention, with a matching lower bound showing no exact…
Read the full surrounding passage
FlashAttention converted that observation into wall-clock speed by combining tiling with recomputation. Blocks of the inputs are loaded into on-chip SRAM, the softmax reduction is performed incrementally across blocks using the online normaliser, and the output is written back without the score matrix ever reaching main memory; for the backward pass, the softmax normalisation factors are kept from the forward pass so that attention can be recomputed on-chip, which the authors state is faster than reading the stored intermediate back from HBM. Their analysis gives O( ) HBM accesses against for standard attention, with a matching lower bound showing no exact algorithm can do asymptotically better across all SRAM sizes [ 1 ] . The reported results were a 7.6 times speedup on the attention computation itself for GPT-2, 15 per cent faster end-to-end BERT-large training against the MLPerf 1.1 record, 3 times on GPT-2 and 2.4 times on long-range arena, plus qualitatively new capability at length: 61.4 per cent on Path-X at 16K and 63.1 per cent on Path-256 at 64K, both above chance for the first time.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to Locality Is the Whole Game: The Memory Hierarchy and What a Kernel Does Not Read