← All parts of this equation

Equation 15 · Part 3 · How Llama's Architecture Actually Works, Generation by Generation

=

G∈{1,2,…,H},G=H⇒multi-head,G=1⇒multi-query,1<G<H⇒grouped-query.G \in \{1, 2, \dots, H\}, \qquad G = H \Rightarrow \text{multi-head}, \quad G = 1 \Rightarrow \text{multi-query}, \quad 1 < G < H \Rightarrow \text{grouped-query}.
=

What this part means

The expressions on both sides represent the same quantity under the stated assumptions.

Its job in the formula

The equals sign connects the complete expression on the left with the complete expression on the right. Both sides must have compatible units.

The passage around this formula

Multi-head attention gives every query head its own key and value projections. That is expensive at inference time for an autoregressive model, because every generated token requires reading the cached keys and values for every previous token, once per head, from memory — and memory bandwidth, not arithmetic, is usually the bottleneck in decoding. Multi-query attention collapses all query heads onto a single shared key-value pair, cutting that cache to a fraction of its multi-head size at some cost in quality. Ainslie and colleagues proposed the middle path that all three Llama generations actually use: grouped-query attention, in which H query heads are partitioned into G groups, and every…

Read this part in the article →

Learn the underlying idea

An equals sign says that the expression on its left and the expression on its right have the same value under the stated definitions and assumptions.

Open the illustrated equality: what the equals sign claims guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.