← Mathematical compendium

Published equation contexts

G∈{1,2,…,H},G=H⇒multi-head,G=1⇒multi-query,1<G<H⇒grouped-queryG \in \{1, 2, \dots, H\}, \qquad G = H \Rightarrow \text{multi-head}, \quad G = 1 \Rightarrow \text{multi-query}, \quad 1 < G < H \Rightarrow \text{grouped-query}

Why this formula appears here

Multi-head attention gives every query head its own key and value projections. That is expensive at inference time for an autoregressive model, because every generated token requires reading the cached keys and values for every previous token, once per head, from memory — and memory bandwidth, not arithmetic, is usually the bottleneck in decoding. Multi-query attention collapses all query heads onto a single shared key-value pair, cutting that cache to a fraction of its multi-head size at some cost in quality. Ainslie and colleagues proposed the middle path that all three Llama generations actually use: grouped-query attention, in which H query heads are partitioned into G groups, and every…

Read the full article-specific guide →

Read the representative guide

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

G∈{1,2,…,H},G=H⇒multi-head,G=1⇒multi-query,1<G<H⇒grouped-query.G \in \{1, 2, \dots, H\}, \qquad G = H \Rightarrow \text{multi-head}, \quad G = 1 \Rightarrow \text{multi-query}, \quad 1 < G < H \Rightarrow \text{grouped-query}.

Equation 15 · Open Models

How Llama's Architecture Actually Works, Generation by Generation

This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions.

Multi-head attention gives every query head its own key and value projections. That is expensive at inference time for an autoregressive model, because every generated token requires reading the cached keys and values for every previous token, once per head, from memory — and memory bandwidth, not arithmetic, is usually the bottleneck in decoding. Multi-query attention collapses all query heads onto a single shared key-value pair, cutting that cache to a fraction of its multi-head size at some cost in quality. Ainslie and colleagues proposed the middle path that all three Llama generations actually use: grouped-query attention, in which H query heads are partitioned into G groups, and every…

Equation guide → · Article →