← All parts of this equation

Equation 15 · Part 1 · How Llama's Architecture Actually Works, Generation by Generation

Symbol G

G∈{1,2,…,H},G=H⇒multi-head,G=1⇒multi-query,1<G<H⇒grouped-query.G \in \{1, 2, \dots, H\}, \qquad G = H \Rightarrow \text{multi-head}, \quad G = 1 \Rightarrow \text{multi-query}, \quad 1 < G < H \Rightarrow \text{grouped-query}.
GG

What this part means

G is part of the quantity the equation computes from the expression on the right.

Its job in the formula

G is part of the quantity the equation computes from the expression on the right.

The passage around this formula

…fraction of its multi-head size at some cost in quality. Ainslie and colleagues proposed the middle path that all three Llama generations actually use: grouped-query attention, in which H query heads are partitioned into G groups, and every query head within a group shares one key-value head pair [ 8 ] . The two earlier schemes are the boundary cases of the same construction: G∈{1,2,…,H},G=H⇒multi-head,G=1⇒multi-query,1<G<H⇒grouped-queryG \in \{1, 2, \dots, H\}, \qquad G = H \Rightarrow \text{multi-head}, \quad G = 1 \Rightarrow \text{multi-query}, \quad 1 < G < H \Rightarrow \text{grouped-query}. Ainslie and colleagues’ own contribution was as much practical as architectural: they showed an existing multi-head…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.