← Back to article

Equation 1 · How AI Memory Systems and the Bandwidth Wall Actually Work

What does this equation mean?

I=FLOPs performedbytes moved,Pachievable=min⁡(Ppeak compute,  I×BWpeak memory)I = \frac{\text{FLOPs performed}}{\text{bytes moved}}, \qquad P_{\text{achievable}} = \min\left(P_{\text{peak compute}},\; I \times BW_{\text{peak memory}}\right)

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Start withFLOPs performed
Divide bybytes moved
This relates toI
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

II

Symbol I

I is the quantity selected or evaluated by the optimization written on the right.

Understand this part →

PachievableP_{\text{achievable}}

Symbol P_achievable

PaP_achievable appears in the objective or constraint used by the optimization on the right.

Understand this part →

Ppeak computeP_{\text{peak compute}}

Symbol P_peak compute

PpP_peak compute appears in the objective or constraint used by the optimization on the right.

Understand this part →

BB

Symbol B

B appears in the objective or constraint used by the optimization on the right.

Understand this part →

Wpeak memoryW_{\text{peak memory}}

Symbol W_peak memory

WpW_peak memory appears in the objective or constraint used by the optimization on the right.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
multiplication

multiplication

Multiply the quantities on either side.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

FLOPs performed\text{FLOPs performed}

Numerator: FLOPs performed

The complete quantity above the fraction bar. FLOPs count floating-point arithmetic operations.

Understand this part →

bytes moved\text{bytes moved}

Denominator: bytes moved

The complete quantity below the fraction bar; it must be nonzero for this division. Bytes measure the data moved or stored.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

The roofline model, introduced by Williams, Waterman, and Patterson in 2009, gives the concept its formal shape. It plots two independent ceilings on the same chart: the chip’s peak arithmetic throughput (operations per second) on one axis, and its peak memory bandwidth (bytes per second) combined with a workload-specific ratio on the other. That ratio — arithmetic intensity — is defined as the number of floating-point operations performed per byte of data moved between memory and the arithmetic units [ 4 ] . A workload’s achievable throughput is bounded by whichever ceiling it hits first: I=FLOPs performedbytes moved,Pachievable=min⁡(Ppeak compute,  I×BWpeak memory)I = \frac{\text{FLOPs performed}}{\text{bytes moved}}, \qquad P_{\text{achievable}} = \min\left(P_{\text{peak compute}},\; I \times BW_{\text{peak memory}}\right). When a workload’s intensity I is high enough that I ×\times BWpeak memoryW_{\text{peak memory}}…
Read the full surrounding passage
The roofline model, introduced by Williams, Waterman, and Patterson in 2009, gives the concept its formal shape. It plots two independent ceilings on the same chart: the chip’s peak arithmetic throughput (operations per second) on one axis, and its peak memory bandwidth (bytes per second) combined with a workload-specific ratio on the other. That ratio — arithmetic intensity — is defined as the number of floating-point operations performed per byte of data moved between memory and the arithmetic units [ 4 ] . A workload’s achievable throughput is bounded by whichever ceiling it hits first: I=FLOPs performedbytes moved,Pachievable=min⁡(Ppeak compute,  I×BWpeak memory)I = \frac{\text{FLOPs performed}}{\text{bytes moved}}, \qquad P_{\text{achievable}} = \min\left(P_{\text{peak compute}},\; I \times BW_{\text{peak memory}}\right). When a workload’s intensity I is high enough that I ×\times BWpeak memoryW_{\text{peak memory}} exceeds the chip’s peak compute rate, the workload is compute-bound: more arithmetic throughput would speed it up, more memory bandwidth would not. When I is low, the workload is memory-bound: the chip’s arithmetic units sit idle waiting for bytes, and only additional bandwidth (or a higher-intensity reformulation of the same computation) helps [ 4 ] . This is the exact mechanism behind the bandwidth-meter-pinned, compute-gauge-idle image below: the two gauges are reading two different ceilings, and only one of them is being pushed against.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to How AI Memory Systems and the Bandwidth Wall Actually Work

See this formula across 1 published context →

Browse the mathematical compendium →