← Mathematical compendium

Published equation contexts

Ω(Nd+N2)\Omega(Nd + N^2)

Why this formula appears here

FlashAttention converted that observation into wall-clock speed by combining tiling with recomputation. Blocks of the inputs are loaded into on-chip SRAM, the softmax reduction is performed incrementally across blocks using the online normaliser, and the output is written back without the score matrix ever reaching main memory; for the backward pass, the softmax normalisation factors are kept from the forward pass so that attention can be recomputed on-chip, which the authors state is faster than reading the stored intermediate back from HBM. Their analysis gives O(N2N^2 d2d^2 M−1M^{-1}) HBM accesses against Ω(Nd+N2)\Omega(Nd + N^2) for standard attention, with a matching lower bound showing no exact…

Read the full article-specific guide →

Read the representative guide

Ω\Omega

Symbol Omega

Omega is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
NN

Symbol N

N is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
dd

Symbol d

d is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Ω(Nd+N2)\Omega(Nd + N^2)

Equation 29 · AI Hardware & Semiconductors

Locality Is the Whole Game: The Memory Hierarchy and What a Kernel Does Not Read

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

FlashAttention converted that observation into wall-clock speed by combining tiling with recomputation. Blocks of the inputs are loaded into on-chip SRAM, the softmax reduction is performed incrementally across blocks using the online normaliser, and the output is written back without the score matrix ever reaching main memory; for the backward pass, the softmax normalisation factors are kept from the forward pass so that attention can be recomputed on-chip, which the authors state is faster than reading the stored intermediate back from HBM. Their analysis gives O(N2N^2 d2d^2 M−1M^{-1}) HBM accesses against Ω(Nd+N2)\Omega(Nd + N^2) for standard attention, with a matching lower bound showing no exact…

Equation guide → · Article →