← All parts of this equation

Equation 19 · Part 5 · How Llama's Architecture Actually Works, Generation by Generation

Symbol d

RMSNorm(x)j=xjRMS(x) gj,RMS(x)=1d∑k=1dxk2+ϵ.\mathrm{RMSNorm}(x)_j = \frac{x_j}{\mathrm{RMS}(x)}\, g_j, \qquad \mathrm{RMS}(x) = \sqrt{\frac{1}{d}\sum_{k=1}^{d} x_k^2 + \epsilon}.
dd

What this part means

d occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Its job in the formula

d occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

The passage around this formula

RMSNorm, from Zhang and Sennrich, made a narrower and more surgical change to normalization. Standard LayerNorm re-centres a layer’s inputs to zero mean and rescales to unit variance before applying a learned gain. Zhang and Sennrich’s hypothesis, tested empirically, was that the re-centring step is not doing useful work — that rescaling invariance alone accounts for LayerNorm’s benefit — and their RMSNorm drops the mean-subtraction entirely: RMSNorm(x)j=xjRMS(x) gj,RMS(x)=1d∑k=1dxk2+ϵ\mathrm{RMSNorm}(x)_j = \frac{x_j}{\mathrm{RMS}(x)}\, g_j, \qquad \mathrm{RMS}(x) = \sqrt{\frac{1}{d}\sum_{k=1}^{d} x_k^2 + \epsilon}. Their reported result was performance comparable to LayerNorm at a running-time reduction of roughly seven to sixty-four percent depending on the model, purely from removing the mean and its gradient computation [ 7 ] . Llama 2’s…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.