← Back to article

Equation 1 · A Rising FLOPs-per-Byte Ratio Explains Why Nvidia Split the Chip in Two

What does this equation mean?

RnR_n

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

the call it. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

RnR_n

Symbol R_n

the call it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Call it RnR_n : the FLOPs-per-second a GPU’s tensor cores can deliver, divided by the bytes-per-second its memory subsystem can supply. The unit that falls out is FLOPs per byte — the arithmetic intensity a workload needs in order to keep that generation’s compute fed rather than starved. This is not a new concept; it is the roofline model, decades old in computer architecture, and the general case for reading any accelerator through it — where the “memory wall” sits, why arithmetic intensity rather than raw FLOPs decides real-world throughput — is already made in full in this publication’s own accelerator architecture explainer , and this piece leans on that framework rather than re-deriving…
Read the full surrounding passage
Call it RnR_n : the FLOPs-per-second a GPU’s tensor cores can deliver, divided by the bytes-per-second its memory subsystem can supply. The unit that falls out is FLOPs per byte — the arithmetic intensity a workload needs in order to keep that generation’s compute fed rather than starved. This is not a new concept; it is the roofline model, decades old in computer architecture, and the general case for reading any accelerator through it — where the “memory wall” sits, why arithmetic intensity rather than raw FLOPs decides real-world throughput — is already made in full in this publication’s own accelerator architecture explainer , and this piece leans on that framework rather than re-deriving it. What has not been done, at least not as a named, generation-by-generation series tied to one vendor’s own product decisions, is compute RnR_n for Nvidia specifically, across enough generations to see whether it is trending anywhere, and then ask what a specific climbing trend would force a chip designer to do about it.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to A Rising FLOPs-per-Byte Ratio Explains Why Nvidia Split the Chip in Two

Browse the mathematical compendium →