← All parts of this equation

Equation 28 · Part 4 · Locality Is the Whole Game: The Memory Hierarchy and What a Kernel Does Not Read

Symbol M^-1

O(N2d2M−1)O(N^2 d^2 M^{-1})
M−1M^{-1}

What this part means

M−M^-1 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Its job in the formula

M−M^-1 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

The passage around this formula

…normalisation factors are kept from the forward pass so that attention can be recomputed on-chip, which the authors state is faster than reading the stored intermediate back from HBM. Their analysis gives O(N2N^2 d2d^2 M−1M^{-1}) HBM accesses against Ω(Nd+N2)\Omega(Nd + N^2) for standard attention, with a matching lower bound showing no exact algorithm can do asymptotically better across all SRAM sizes [ 1 ] . The reported results were a 7.6 times speedup on the attention computation itself for GPT-2, 15 per cent faster…

Read this part in the article →

Learn the underlying idea

An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.

Open the illustrated exponents: repeated multiplication and powers guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.