← Mathematical compendium

Published equation contexts

I=FLOPs performedbytes read and writtenI = \frac{\text{FLOPs performed}}{\text{bytes read and written}}

Why this formula appears here

The standard way to make this precise is a piece of hardware-performance reasoning called the roofline model, and stating it once explains why every mitigation below targets one phase and not the other. Every accelerator has two ceilings: a peak arithmetic rate, in floating-point operations per second, and a peak memory-bandwidth rate, in bytes per second. Which one binds a given piece of work depends on that work’s arithmetic intensity — the ratio of arithmetic performed to bytes moved: I=FLOPs performedbytes read and writtenI = \frac{\text{FLOPs performed}}{\text{bytes read and written}}. A computation is compute-bound when I exceeds the hardware’s own ratio of FLOPs to bandwidth, and memory-bound when it falls below it. Prefill, processing every position of a prompt…

Read the full article-specific guide →

Read the representative guide

FLOPs performed\text{FLOPs performed}

Numerator: FLOPs performed

The complete quantity above the fraction bar. FLOPs count floating-point arithmetic operations.

Read this term in its guide →
bytes read and written\text{bytes read and written}

Denominator: bytes read and written

The complete quantity below the fraction bar; it must be nonzero for this division. Bytes measure the data moved or stored.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

I=FLOPs performedbytes read and writtenI = \frac{\text{FLOPs performed}}{\text{bytes read and written}}

Equation 3 · Inference Economics

How AI Inference Serving Actually Works

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

The standard way to make this precise is a piece of hardware-performance reasoning called the roofline model, and stating it once explains why every mitigation below targets one phase and not the other. Every accelerator has two ceilings: a peak arithmetic rate, in floating-point operations per second, and a peak memory-bandwidth rate, in bytes per second. Which one binds a given piece of work depends on that work’s arithmetic intensity — the ratio of arithmetic performed to bytes moved: I=FLOPs performedbytes read and writtenI = \frac{\text{FLOPs performed}}{\text{bytes read and written}}. A computation is compute-bound when I exceeds the hardware’s own ratio of FLOPs to bandwidth, and memory-bound when it falls below it. Prefill, processing every position of a prompt…

Meanings in this article

  • II: the ratio of arithmetic performed to bytes moved.
Equation guide → · Article →