← Back to article

Equation 10 · How AI Accelerator Architecture Actually Works

What does this equation mean?

T≈(r+c−2)+L,T \approx (r + c - 2) + L,

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

TT

Symbol T

T is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

rr

Symbol r

r is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

cc

Symbol c

c is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

LL

Symbol L

the number of positions in each sequence.

Understand this part →

≈

≈

Approximately equal to; the equality is not exact.

Understand this part →

addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subtraction

subtraction

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Understand this part →

How to interpret it

Its accuracy depends on the assumptions and range of use described in the article.

What the article says around this equation

Consider an idealised r ×\times c array holding one operand stationary, with a second operand streamed through in a sequence of length L — for instance, L activation vectors passed through a weight-stationary array of r input rows and c output columns. Before the first result emerges, the data has to propagate across the array’s diagonal, costing roughly r + c - 2 cycles; after the last input enters, the same number of cycles is needed to drain the pipeline. Total latency is therefore approximately T≈(r+c−2)+LT \approx (r + c - 2) + L. while the useful work performed is r ⋅\cdot c ⋅\cdot L multiply-accumulates — one per cell, once per streamed step, in steady state. Dividing useful work by the total cell-cycles…
Read the full surrounding passage
Consider an idealised r ×\times c array holding one operand stationary, with a second operand streamed through in a sequence of length L — for instance, L activation vectors passed through a weight-stationary array of r input rows and c output columns. Before the first result emerges, the data has to propagate across the array’s diagonal, costing roughly r + c - 2 cycles; after the last input enters, the same number of cycles is needed to drain the pipeline. Total latency is therefore approximately T≈(r+c−2)+LT \approx (r + c - 2) + L. while the useful work performed is r ⋅\cdot c ⋅\cdot L multiply-accumulates — one per cell, once per streamed step, in steady state. Dividing useful work by the total cell-cycles available, r ⋅\cdot c ⋅\cdot T , gives an idealised utilization

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to How AI Accelerator Architecture Actually Works

See this formula across 1 published context →

Browse the mathematical compendium →