← Mathematical compendium

Published equation contexts

RMSNorm(x)j=xjRMS(x) gj,RMS(x)=1d∑k=1dxk2+ϵ\mathrm{RMSNorm}(x)_j = \frac{x_j}{\mathrm{RMS}(x)}\, g_j, \qquad \mathrm{RMS}(x) = \sqrt{\frac{1}{d}\sum_{k=1}^{d} x_k^2 + \epsilon}

Why this formula appears here

RMSNorm, from Zhang and Sennrich, made a narrower and more surgical change to normalization. Standard LayerNorm re-centres a layer’s inputs to zero mean and rescales to unit variance before applying a learned gain. Zhang and Sennrich’s hypothesis, tested empirically, was that the re-centring step is not doing useful work — that rescaling invariance alone accounts for LayerNorm’s benefit — and their RMSNorm drops the mean-subtraction entirely: RMSNorm(x)j=xjRMS(x) gj,RMS(x)=1d∑k=1dxk2+ϵ\mathrm{RMSNorm}(x)_j = \frac{x_j}{\mathrm{RMS}(x)}\, g_j, \qquad \mathrm{RMS}(x) = \sqrt{\frac{1}{d}\sum_{k=1}^{d} x_k^2 + \epsilon}. Their reported result was performance comparable to LayerNorm at a running-time reduction of roughly seven to sixty-four percent depending on the model, purely from removing the mean and its gradient computation [ 7 ] . Llama 2’s…

Read the full article-specific guide →

Read the representative guide

dd

Symbol d

d occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →
kk

Symbol k

k appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Read this term in its guide →
RMS(x)\mathrm{RMS}(x)

Denominator: RMS(x)

The complete quantity below the fraction bar; it must be nonzero for this division.

Read this term in its guide →
k=1k=1

Starting index or lower bound: k=1

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Read this term in its guide →
dd

Ending index or upper bound: d

This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

RMSNorm(x)j=xjRMS(x) gj,RMS(x)=1d∑k=1dxk2+ϵ.\mathrm{RMSNorm}(x)_j = \frac{x_j}{\mathrm{RMS}(x)}\, g_j, \qquad \mathrm{RMS}(x) = \sqrt{\frac{1}{d}\sum_{k=1}^{d} x_k^2 + \epsilon}.

Equation 19 · Open Models

How Llama's Architecture Actually Works, Generation by Generation

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

RMSNorm, from Zhang and Sennrich, made a narrower and more surgical change to normalization. Standard LayerNorm re-centres a layer’s inputs to zero mean and rescales to unit variance before applying a learned gain. Zhang and Sennrich’s hypothesis, tested empirically, was that the re-centring step is not doing useful work — that rescaling invariance alone accounts for LayerNorm’s benefit — and their RMSNorm drops the mean-subtraction entirely: RMSNorm(x)j=xjRMS(x) gj,RMS(x)=1d∑k=1dxk2+ϵ\mathrm{RMSNorm}(x)_j = \frac{x_j}{\mathrm{RMS}(x)}\, g_j, \qquad \mathrm{RMS}(x) = \sqrt{\frac{1}{d}\sum_{k=1}^{d} x_k^2 + \epsilon}. Their reported result was performance comparable to LayerNorm at a running-time reduction of roughly seven to sixty-four percent depending on the model, purely from removing the mean and its gradient computation [ 7 ] . Llama 2’s…

Equation guide → · Article →