← Mathematical compendium

Published equation contexts

P  ≤  min⁡ ⁣(Pmax⁡,  I⋅B),I=FLOPsbytes movedP \;\le\; \min\!\left(P_{\max},\; I \cdot B\right), \qquad I = \frac{\text{FLOPs}}{\text{bytes moved}}

Why this formula appears here

Given a memory system built this way, the question for any specific piece of computation is simple to state and consequential to answer: for this kernel, is the bottleneck the arithmetic units or the memory system that feeds them? The quantity that answers it is arithmetic intensity , defined as the number of floating-point operations a kernel performs per byte it moves across the memory boundary that matters. Williams, Waterman and Patterson formalized the relationship between intensity and attainable performance as the roofline model: P  ≤  min⁡ ⁣(Pmax⁡,  I⋅B),I=FLOPsbytes movedP \;\le\; \min\!\left(P_{\max},\; I \cdot B\right), \qquad I = \frac{\text{FLOPs}}{\text{bytes moved}}. with Pmax⁡P_{\max} the peak arithmetic rate of the device, B the achievable bandwidth of the memory tier supplying the operands, and P the…

Read the full article-specific guide →

Read the representative guide

FLOPs\text{FLOPs}

Numerator: FLOPs

The complete quantity above the fraction bar. FLOPs count floating-point arithmetic operations.

Read this term in its guide →
bytes moved\text{bytes moved}

Denominator: bytes moved

The complete quantity below the fraction bar; it must be nonzero for this division. Bytes measure the data moved or stored.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

P  ≤  min⁡ ⁣(Pmax⁡,  I⋅B),I=FLOPsbytes movedP \;\le\; \min\!\left(P_{\max},\; I \cdot B\right), \qquad I = \frac{\text{FLOPs}}{\text{bytes moved}}

Equation 1 · Semiconductors

How AI Memory Systems and the Bandwidth Wall Actually Work

This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions.

Given a memory system built this way, the question for any specific piece of computation is simple to state and consequential to answer: for this kernel, is the bottleneck the arithmetic units or the memory system that feeds them? The quantity that answers it is arithmetic intensity , defined as the number of floating-point operations a kernel performs per byte it moves across the memory boundary that matters. Williams, Waterman and Patterson formalized the relationship between intensity and attainable performance as the roofline model: P  ≤  min⁡ ⁣(Pmax⁡,  I⋅B),I=FLOPsbytes movedP \;\le\; \min\!\left(P_{\max},\; I \cdot B\right), \qquad I = \frac{\text{FLOPs}}{\text{bytes moved}}. with Pmax⁡P_{\max} the peak arithmetic rate of the device, B the achievable bandwidth of the memory tier supplying the operands, and P the…

Meanings in this article

  • PP: the attainable performance [ 11 ].
  • Pmax⁡P_{\max}: the peak arithmetic rate of the device.
  • BB: the achievable bandwidth of the memory tier supplying the operands.
Equation guide → · Article →