Symbol G
G is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
Multi-head attention gives every query head its own key and value projections. That is expensive at inference time for an autoregressive model, because every generated token requires reading the cached keys and values for every previous token, once per head, from memory — and memory bandwidth, not arithmetic, is usually the bottleneck in decoding. Multi-query attention collapses all query heads onto a single shared key-value pair, cutting that cache to a fraction of its multi-head size at some cost in quality. Ainslie and colleagues proposed the middle path that all three Llama generations actually use: grouped-query attention, in which H query heads are partitioned into G groups, and every…
G is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →H is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 15 · Open Models
This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions.
Multi-head attention gives every query head its own key and value projections. That is expensive at inference time for an autoregressive model, because every generated token requires reading the cached keys and values for every previous token, once per head, from memory — and memory bandwidth, not arithmetic, is usually the bottleneck in decoding. Multi-query attention collapses all query heads onto a single shared key-value pair, cutting that cache to a fraction of its multi-head size at some cost in quality. Ainslie and colleagues proposed the middle path that all three Llama generations actually use: grouped-query attention, in which H query heads are partitioned into G groups, and every…
Equation guide → · Article →