Equation 2 · A Rising FLOPs-per-Byte Ratio Explains Why Nvidia Split the Chip in Two
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the call it. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Call it : the FLOPs-per-second a GPU’s tensor cores can deliver, divided by the bytes-per-second its memory subsystem can supply. The unit that falls out is FLOPs per byte — the arithmetic intensity a workload needs in order to keep that generation’s compute fed rather than starved. This is not a new concept; it is the roofline model, decades old in computer architecture, and the general case for reading any accelerator through it — where the “memory wall” sits, why arithmetic intensity rather than raw FLOPs decides real-world throughput — is already made in full in this publication’s own accelerator architecture explainer , and this piece leans on that framework rather than re-deriving…
Read the full surrounding passage
Call it : the FLOPs-per-second a GPU’s tensor cores can deliver, divided by the bytes-per-second its memory subsystem can supply. The unit that falls out is FLOPs per byte — the arithmetic intensity a workload needs in order to keep that generation’s compute fed rather than starved. This is not a new concept; it is the roofline model, decades old in computer architecture, and the general case for reading any accelerator through it — where the “memory wall” sits, why arithmetic intensity rather than raw FLOPs decides real-world throughput — is already made in full in this publication’s own accelerator architecture explainer , and this piece leans on that framework rather than re-deriving it. What has not been done, at least not as a named, generation-by-generation series tied to one vendor’s own product decisions, is compute for Nvidia specifically, across enough generations to see whether it is trending anywhere, and then ask what a specific climbing trend would force a chip designer to do about it.
Sources cited in the article section
- [4] NVIDIA H100 Tensor Core GPU Datasheet ↗
- [5] Hopper (microarchitecture) ↗
- [6] NVIDIA HGX Platform (B200 specifications) ↗
These citations give research context. Read each source to check which claims it supports.
Return to A Rising FLOPs-per-Byte Ratio Explains Why Nvidia Split the Chip in Two