← All parts of this equation

Equation 18 · Part 14 · How a Model Actually Gets Small Enough to Run on a Phone

superscript

W′=arg min⁡∥W′∥0≤k∥WX−W′X∥22W' = \operatorname*{arg\,min}_{\|W'\|_0 \le k} \|WX - W'X\|_2^2
superscript

What this part means

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Its job in the formula

A raised mark can be a power or an index. Its position and the surrounding notation determine which.

The passage around this formula

SparseGPT poses pruning as a per-layer reconstruction problem. For a layer with weight matrix W and a small calibration set of activations X , it looks for a sparse replacement W' that keeps that layer’s output as close as possible to the original: W′=arg min⁡∥W′∥0≤k∥WX−W′X∥22W' = \operatorname*{arg\,min}_{\|W'\|_0 \le k} \|WX - W'X\|_2^2. Rather than solving this by retraining, Frantar and Alistarh adapt a closed-form update derived from the layer’s second-order (Hessian) information, in the spirit of the older Optimal Brain Surgeon method, so that whenever a weight is removed the remaining weights in that row are analytically nudged to compensate for its absence. The result, reported for the GPT-family models tested, is that “large-scale generative pretrained…

Read this part in the article →

Learn the underlying idea

An exponent tells how a base is used in multiplication. In x³, x is the base and 3 is the exponent: x³ = x × x × x.

Open the illustrated exponents: repeated multiplication and powers guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.