Symbol Δ W
Δ W is the quantity selected or evaluated by the optimization written on the right.
Read this term in its guide →Published equation contexts
When the adapter route is chosen, the engineering default within it is equally clear: adapt, do not retrain. Low-Rank Adaptation freezes the pretrained weights and injects a pair of small trainable matrices into selected layers, so that a weight update is expressed as a low-rank product rather than a dense matrix the size of the original layer, . with the forward pass computing h = x + W x against the frozen base weight . The method’s authors report reducing the number of trainable parameters by roughly ten thousand times and GPU memory requirements by roughly three times relative to full fine-tuning of a 175-billion-parameter model, while matching or…
Δ W is the quantity selected or evaluated by the optimization written on the right.
Read this term in its guide →B appears in the objective or constraint used by the optimization on the right.
Read this term in its guide →A appears in the objective or constraint used by the optimization on the right.
Read this term in its guide →× r appears in the objective or constraint used by the optimization on the right.
Read this term in its guide →× k appears in the objective or constraint used by the optimization on the right.
Read this term in its guide →r appears in the objective or constraint used by the optimization on the right.
Read this term in its guide →d appears in the objective or constraint used by the optimization on the right.
Read this term in its guide →k appears in the objective or constraint used by the optimization on the right.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 1 · Foundation Models
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
When the adapter route is chosen, the engineering default within it is equally clear: adapt, do not retrain. Low-Rank Adaptation freezes the pretrained weights and injects a pair of small trainable matrices into selected layers, so that a weight update is expressed as a low-rank product rather than a dense matrix the size of the original layer, . with the forward pass computing h = x + W x against the frozen base weight . The method’s authors report reducing the number of trainable parameters by roughly ten thousand times and GPU memory requirements by roughly three times relative to full fine-tuning of a 175-billion-parameter model, while matching or…
Equation guide → · Article →