← Back to article

Equation 29 · Locality Is the Whole Game: The Memory Hierarchy and What a Kernel Does Not Read

What does this equation mean?

Ω(Nd+N2)\Omega(Nd + N^2)

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

Ω\Omega

Symbol Omega

Omega is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

NN

Symbol N

N is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

dd

Symbol d

d is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

N2N^2

Symbol N^2

The square of N: multiply N by itself.

Understand this part →

addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

FlashAttention converted that observation into wall-clock speed by combining tiling with recomputation. Blocks of the inputs are loaded into on-chip SRAM, the softmax reduction is performed incrementally across blocks using the online normaliser, and the output is written back without the score matrix ever reaching main memory; for the backward pass, the softmax normalisation factors are kept from the forward pass so that attention can be recomputed on-chip, which the authors state is faster than reading the stored intermediate back from HBM. Their analysis gives O(N2N^2 d2d^2 M−1M^{-1}) HBM accesses against Ω(Nd+N2)\Omega(Nd + N^2) for standard attention, with a matching lower bound showing no exact…
Read the full surrounding passage
FlashAttention converted that observation into wall-clock speed by combining tiling with recomputation. Blocks of the inputs are loaded into on-chip SRAM, the softmax reduction is performed incrementally across blocks using the online normaliser, and the output is written back without the score matrix ever reaching main memory; for the backward pass, the softmax normalisation factors are kept from the forward pass so that attention can be recomputed on-chip, which the authors state is faster than reading the stored intermediate back from HBM. Their analysis gives O(N2N^2 d2d^2 M−1M^{-1}) HBM accesses against Ω(Nd+N2)\Omega(Nd + N^2) for standard attention, with a matching lower bound showing no exact algorithm can do asymptotically better across all SRAM sizes [ 1 ] . The reported results were a 7.6 times speedup on the attention computation itself for GPT-2, 15 per cent faster end-to-end BERT-large training against the MLPerf 1.1 record, 3 times on GPT-2 and 2.4 times on long-range arena, plus qualitatively new capability at length: 61.4 per cent on Path-X at 16K and 63.1 per cent on Path-256 at 64K, both above chance for the first time.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Locality Is the Whole Game: The Memory Hierarchy and What a Kernel Does Not Read

See this formula across 1 published context →

Browse the mathematical compendium →