← Back to article

Equation 19 · How Llama's Architecture Actually Works, Generation by Generation

What does this equation mean?

RMSNorm(x)j=xjRMS(x) gj,RMS(x)=1d∑k=1dxk2+ϵ.\mathrm{RMSNorm}(x)_j = \frac{x_j}{\mathrm{RMS}(x)}\, g_j, \qquad \mathrm{RMS}(x) = \sqrt{\frac{1}{d}\sum_{k=1}^{d} x_k^2 + \epsilon}.

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Start withx_j
Divide byRMS(x)
This relates toRMSNorm(x)_j
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

xx

Symbol x

x is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

jj

Symbol j

j occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

xjx_j

Symbol x_j

xjx_j occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Understand this part →

gjg_j

Symbol g_j

gjg_j is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

dd

Symbol d

d occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Understand this part →

kk

Symbol k

k appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Understand this part →

xk2x_k^2

Symbol x_k^2

xk2x_k^2 is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

ϵ\epsilon

Symbol epsilon

epsilon is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
√

√

Take a square root.

Understand this part →

addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →
RMS(x)\mathrm{RMS}(x)

Denominator: RMS(x)

The complete quantity below the fraction bar; it must be nonzero for this division.

Understand this part →

11

Numerator: 1

The complete quantity above the fraction bar.

Understand this part →

k=1k=1

Starting index or lower bound: k=1

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Understand this part →

dd

Ending index or upper bound: d

This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

RMSNorm, from Zhang and Sennrich, made a narrower and more surgical change to normalization. Standard LayerNorm re-centres a layer’s inputs to zero mean and rescales to unit variance before applying a learned gain. Zhang and Sennrich’s hypothesis, tested empirically, was that the re-centring step is not doing useful work — that rescaling invariance alone accounts for LayerNorm’s benefit — and their RMSNorm drops the mean-subtraction entirely: RMSNorm(x)j=xjRMS(x) gj,RMS(x)=1d∑k=1dxk2+ϵ\mathrm{RMSNorm}(x)_j = \frac{x_j}{\mathrm{RMS}(x)}\, g_j, \qquad \mathrm{RMS}(x) = \sqrt{\frac{1}{d}\sum_{k=1}^{d} x_k^2 + \epsilon}. Their reported result was performance comparable to LayerNorm at a running-time reduction of roughly seven to sixty-four percent depending on the model, purely from removing the mean and its gradient computation [ 7 ] . Llama 2’s…
Read the full surrounding passage
RMSNorm, from Zhang and Sennrich, made a narrower and more surgical change to normalization. Standard LayerNorm re-centres a layer’s inputs to zero mean and rescales to unit variance before applying a learned gain. Zhang and Sennrich’s hypothesis, tested empirically, was that the re-centring step is not doing useful work — that rescaling invariance alone accounts for LayerNorm’s benefit — and their RMSNorm drops the mean-subtraction entirely: RMSNorm(x)j=xjRMS(x) gj,RMS(x)=1d∑k=1dxk2+ϵ\mathrm{RMSNorm}(x)_j = \frac{x_j}{\mathrm{RMS}(x)}\, g_j, \qquad \mathrm{RMS}(x) = \sqrt{\frac{1}{d}\sum_{k=1}^{d} x_k^2 + \epsilon}. Their reported result was performance comparable to LayerNorm at a running-time reduction of roughly seven to sixty-four percent depending on the model, purely from removing the mean and its gradient computation [ 7 ] . Llama 2’s architecture section lists RMSNorm as the pre-normalization scheme applied before each attention and feed-forward sublayer, and it has not been revisited in any subsequent generation’s public documentation [ 1 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to How Llama's Architecture Actually Works, Generation by Generation

See this formula across 1 published context →

Browse the mathematical compendium →