← Back to article

Equation 18 · How a Model Actually Gets Small Enough to Run on a Phone

What does this equation mean?

W′=arg min⁡∥W′∥0≤k∥WX−W′X∥22W' = \operatorname*{arg\,min}_{\|W'\|_0 \le k} \|WX - W'X\|_2^2

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationsoperatorname*argmin_|W'|_0 ≤ k |WX - W'X|_2^2
Result or conditionW'
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

WW

Symbol W

W is part of the quantity the equation computes from the expression on the right.

Understand this part →

aa

Symbol a

a is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

rr

Symbol r

r is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

gg

Symbol g

g is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

mm

Symbol m

m is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

ii

Symbol i

i is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

nn

Symbol n

n is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

kk

Symbol k

k is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

XX

Symbol X

X is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
subtraction

subtraction

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

SparseGPT poses pruning as a per-layer reconstruction problem. For a layer with weight matrix W and a small calibration set of activations X , it looks for a sparse replacement W' that keeps that layer’s output as close as possible to the original: W′=arg min⁡∥W′∥0≤k∥WX−W′X∥22W' = \operatorname*{arg\,min}_{\|W'\|_0 \le k} \|WX - W'X\|_2^2. Rather than solving this by retraining, Frantar and Alistarh adapt a closed-form update derived from the layer’s second-order (Hessian) information, in the spirit of the older Optimal Brain Surgeon method, so that whenever a weight is removed the remaining weights in that row are analytically nudged to compensate for its absence. The result, reported for the GPT-family models tested, is that “large-scale generative pretrained…
Read the full surrounding passage
SparseGPT poses pruning as a per-layer reconstruction problem. For a layer with weight matrix W and a small calibration set of activations X , it looks for a sparse replacement W' that keeps that layer’s output as close as possible to the original: W′=arg min⁡∥W′∥0≤k∥WX−W′X∥22W' = \operatorname*{arg\,min}_{\|W'\|_0 \le k} \|WX - W'X\|_2^2. Rather than solving this by retraining, Frantar and Alistarh adapt a closed-form update derived from the layer’s second-order (Hessian) information, in the spirit of the older Optimal Brain Surgeon method, so that whenever a weight is removed the remaining weights in that row are analytically nudged to compensate for its absence. The result, reported for the GPT-family models tested, is that “large-scale generative pretrained transformer family models can be pruned to at least 50% sparsity in one-shot, without any retraining, at minimal loss of accuracy,” with a 175-billion-parameter model prunable in a few hours on a single GPU, and the method extends to semi-structured 2:4 patterns that specific accelerator tensor cores can execute directly [ 5 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to How a Model Actually Gets Small Enough to Run on a Phone

See this formula across 1 published context →

Browse the mathematical compendium →