← Mathematical compendium

Published equation contexts

W′=arg min⁡∥W′∥0≤k∥WX−W′X∥22W' = \operatorname*{arg\,min}_{\|W'\|_0 \le k} \|WX - W'X\|_2^2

Why this formula appears here

SparseGPT poses pruning as a per-layer reconstruction problem. For a layer with weight matrix W and a small calibration set of activations X , it looks for a sparse replacement W' that keeps that layer’s output as close as possible to the original: W′=arg min⁡∥W′∥0≤k∥WX−W′X∥22W' = \operatorname*{arg\,min}_{\|W'\|_0 \le k} \|WX - W'X\|_2^2. Rather than solving this by retraining, Frantar and Alistarh adapt a closed-form update derived from the layer’s second-order (Hessian) information, in the spirit of the older Optimal Brain Surgeon method, so that whenever a weight is removed the remaining weights in that row are analytically nudged to compensate for its absence. The result, reported for the GPT-family models tested, is that “large-scale generative pretrained…

Read the full article-specific guide →

Read the representative guide

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

W′=arg min⁡∥W′∥0≤k∥WX−W′X∥22W' = \operatorname*{arg\,min}_{\|W'\|_0 \le k} \|WX - W'X\|_2^2

Equation 18 · Edge AI & Electronics

How a Model Actually Gets Small Enough to Run on a Phone

This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions.

SparseGPT poses pruning as a per-layer reconstruction problem. For a layer with weight matrix W and a small calibration set of activations X , it looks for a sparse replacement W' that keeps that layer’s output as close as possible to the original: W′=arg min⁡∥W′∥0≤k∥WX−W′X∥22W' = \operatorname*{arg\,min}_{\|W'\|_0 \le k} \|WX - W'X\|_2^2. Rather than solving this by retraining, Frantar and Alistarh adapt a closed-form update derived from the layer’s second-order (Hessian) information, in the spirit of the older Optimal Brain Surgeon method, so that whenever a weight is removed the remaining weights in that row are analytically nudged to compensate for its absence. The result, reported for the GPT-family models tested, is that “large-scale generative pretrained…

Equation guide → · Article →