← Back to article

Equation 1 · Comparing the Main Approaches to AI Accelerator Architecture

What does this equation mean?

Tachieved=Tpeak⋅Uhw⋅UcompilerT_{\mathrm{achieved}} = T_{\mathrm{peak}} \cdot U_{\mathrm{hw}} \cdot U_{\mathrm{compiler}}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationsT_peak × U_hw × U_compiler
Result or conditionT_achieved
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

TachievedT_{\mathrm{achieved}}

Symbol T_achieved

TaT_achieved is part of the quantity the equation computes from the expression on the right.

Understand this part →

TpeakT_{\mathrm{peak}}

Symbol T_peak

TpT_peak is one factor in the product that computes the quantity on the left.

Understand this part →

UhwU_{\mathrm{hw}}

Symbol U_hw

UhU_hw is one factor in the product that computes the quantity on the left.

Understand this part →

UcompilerU_{\mathrm{compiler}}

Symbol U_compiler

UcU_compiler is one factor in the product that computes the quantity on the left.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
multiplication

multiplication

Multiply the quantities on either side.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

It is tempting to compare these four families by a single number — peak operations per second, or operations per watt — and the rest of this article explains why that number, alone, is close to meaningless. A more useful decomposition separates what a device can theoretically do from what a piece of software actually gets it to do, and splits that gap into two distinct causes: Tachieved=Tpeak⋅Uhw⋅UcompilerT_{\mathrm{achieved}} = T_{\mathrm{peak}} \cdot U_{\mathrm{hw}} \cdot U_{\mathrm{compiler}}. Here TpeakT_{\mathrm{peak}} is the arithmetic identity fixed at design time, UhwU_{\mathrm{hw}} ∈\in (0,1] is the fraction of cycles the hardware keeps its arithmetic units genuinely busy on whatever workload is thrown at it, and UcompilerU_{\mathrm{compiler}} ∈\in (0,1] is the fraction of a program’s theoretically…
Read the full surrounding passage
It is tempting to compare these four families by a single number — peak operations per second, or operations per watt — and the rest of this article explains why that number, alone, is close to meaningless. A more useful decomposition separates what a device can theoretically do from what a piece of software actually gets it to do, and splits that gap into two distinct causes: Tachieved=Tpeak⋅Uhw⋅UcompilerT_{\mathrm{achieved}} = T_{\mathrm{peak}} \cdot U_{\mathrm{hw}} \cdot U_{\mathrm{compiler}}. Here TpeakT_{\mathrm{peak}} is the arithmetic identity fixed at design time, UhwU_{\mathrm{hw}} ∈\in (0,1] is the fraction of cycles the hardware keeps its arithmetic units genuinely busy on whatever workload is thrown at it, and UcompilerU_{\mathrm{compiler}} ∈\in (0,1] is the fraction of a program’s theoretically available parallelism that the compiler or mapper actually manages to expose to the hardware. A GPU’s SIMT scheduler mostly targets UhwU_{\mathrm{hw}} , hiding latency and filling gaps dynamically regardless of how well the source program was written. A systolic array and a statically scheduled dataflow chip push almost the entire burden onto UcompilerU_{\mathrm{compiler}} : there is no runtime mechanism left to rescue a poorly mapped program. Sze and colleagues make exactly this point about dataflow choice inside a fixed processing-element array, cataloguing weight-stationary, output-stationary, row-stationary and no-local-reuse dataflows that each keep a different operand fixed in the register file to maximise reuse for a given data-movement energy budget, and observing that because “all of the variables are known before runtime,” an offline mapper can be built to choose the energy-optimal dataflow for a given layer shape and hardware configuration [ 10 ] . That is UcompilerU_{\mathrm{compiler}} as an explicit design target rather than something left to chance.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Comparing the Main Approaches to AI Accelerator Architecture

See this formula across 1 published context →

Browse the mathematical compendium →