Equation 18 · How a Model Actually Gets Small Enough to Run on a Phone
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol W
W is part of the quantity the equation computes from the expression on the right.
Symbol a
a is one of the signed contributions combined to compute the quantity on the left.
Symbol r
r is one of the signed contributions combined to compute the quantity on the left.
Symbol g
g is one of the signed contributions combined to compute the quantity on the left.
Symbol m
m is one of the signed contributions combined to compute the quantity on the left.
Symbol i
i is one of the signed contributions combined to compute the quantity on the left.
Symbol n
n is one of the signed contributions combined to compute the quantity on the left.
Symbol k
k is one of the signed contributions combined to compute the quantity on the left.
Symbol X
X is one of the signed contributions combined to compute the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
SparseGPT poses pruning as a per-layer reconstruction problem. For a layer with weight matrix W and a small calibration set of activations X , it looks for a sparse replacement W' that keeps that layer’s output as close as possible to the original: . Rather than solving this by retraining, Frantar and Alistarh adapt a closed-form update derived from the layer’s second-order (Hessian) information, in the spirit of the older Optimal Brain Surgeon method, so that whenever a weight is removed the remaining weights in that row are analytically nudged to compensate for its absence. The result, reported for the GPT-family models tested, is that “large-scale generative pretrained…
Read the full surrounding passage
SparseGPT poses pruning as a per-layer reconstruction problem. For a layer with weight matrix W and a small calibration set of activations X , it looks for a sparse replacement W' that keeps that layer’s output as close as possible to the original: . Rather than solving this by retraining, Frantar and Alistarh adapt a closed-form update derived from the layer’s second-order (Hessian) information, in the spirit of the older Optimal Brain Surgeon method, so that whenever a weight is removed the remaining weights in that row are analytically nudged to compensate for its absence. The result, reported for the GPT-family models tested, is that “large-scale generative pretrained transformer family models can be pruned to at least 50% sparsity in one-shot, without any retraining, at minimal loss of accuracy,” with a 175-billion-parameter model prunable in a few hours on a single GPU, and the method extends to semi-structured 2:4 patterns that specific accelerator tensor cores can execute directly [ 5 ] .
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to How a Model Actually Gets Small Enough to Run on a Phone