Equation 6 · How AI Accelerator Architecture Actually Works
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the number of positions in each sequence. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Consider an idealised r c array holding one operand stationary, with a second operand streamed through in a sequence of length L — for instance, L activation vectors passed through a weight-stationary array of r input rows and c output columns. Before the first result emerges, the data has to propagate across the array’s diagonal, costing roughly r + c - 2 cycles; after the last input enters, the same number of cycles is needed to drain the pipeline. Total latency is therefore approximately
Sources cited in the article section
- [7] Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore Architectures ↗
- [2] In-Datacenter Performance Analysis of a Tensor Processing Unit ↗
- [6] SCALE-Sim: Systolic CNN Accelerator Simulator ↗
These citations give research context. Read each source to check which claims it supports.