← Mathematical compendium

Published equation contexts

Qmin⁡  ≈  23 N3MQ_{\min} \;\approx\; 2\sqrt{3}\,\frac{N^3}{\sqrt{M}}

Why this formula appears here

Taking b as large as the constraint allows gives Qmin⁡  ≈  23 N3MQ_{\min} \;\approx\; 2\sqrt{3}\,\frac{N^3}{\sqrt{M}} . Two things in that expression matter more than the constant. First, traffic falls linearly in b : doubling the tile halves the bytes moved, which is why tiling is the single highest-leverage transformation in dense linear algebra. Second, and less often noticed, traffic falls only as M−1/2M^{-1/2} . Quadrupling the fast store buys a factor of two in traffic, not a factor of four. That square root is the reason architects cannot simply spend their way out of the problem with more SRAM, and it is worth holding on to when reading any claim that a larger cache will fix a bandwidth-bound kernel.

Read the full article-specific guide →

Read the representative guide

Qmin⁡Q_{\min}

Symbol Q_min

QmQ_min is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
MM

Symbol M

M occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Its accuracy depends on the assumptions and range of use described in the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Qmin⁡  ≈  23 N3M.Q_{\min} \;\approx\; 2\sqrt{3}\,\frac{N^3}{\sqrt{M}} .

Equation 19 · AI Hardware & Semiconductors

Locality Is the Whole Game: The Memory Hierarchy and What a Kernel Does Not Read

This equation gives an approximation: it relates the quantities while allowing an approximation.

Taking b as large as the constraint allows gives Qmin⁡  ≈  23 N3MQ_{\min} \;\approx\; 2\sqrt{3}\,\frac{N^3}{\sqrt{M}} . Two things in that expression matter more than the constant. First, traffic falls linearly in b : doubling the tile halves the bytes moved, which is why tiling is the single highest-leverage transformation in dense linear algebra. Second, and less often noticed, traffic falls only as M−1/2M^{-1/2} . Quadrupling the fast store buys a factor of two in traffic, not a factor of four. That square root is the reason architects cannot simply spend their way out of the problem with more SRAM, and it is worth holding on to when reading any claim that a larger cache will fix a bandwidth-bound kernel.

Equation guide → · Article →