Symbol W
W is the quantity selected or evaluated by the optimization written on the right.
Read this term in its guide →Published equation contexts
Low-Rank Adaptation, LoRA, takes a different approach: it freezes the pretrained weight matrix entirely and represents the update as the product of two much smaller matrices. For a frozen weight matrix , LoRA represents the adapted weight as . Only B and A are trained; never moves during fine-tuning [ 5 ] . The effect on trainable parameter count for that one matrix is to fall from dk to r(d + k) , which is small whenever the chosen rank r is small relative to d and k — and because BA can be merged back into after training, LoRA adds no extra inference latency once deployed. Hu and colleagues report, for their comparison against full…
W is the quantity selected or evaluated by the optimization written on the right.
Read this term in its guide →appears in the objective or constraint used by the optimization on the right.
Read this term in its guide →Δ W appears in the objective or constraint used by the optimization on the right.
Read this term in its guide →× r appears in the objective or constraint used by the optimization on the right.
Read this term in its guide →× k appears in the objective or constraint used by the optimization on the right.
Read this term in its guide →r appears in the objective or constraint used by the optimization on the right.
Read this term in its guide →d appears in the objective or constraint used by the optimization on the right.
Read this term in its guide →k appears in the objective or constraint used by the optimization on the right.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 2 · Open Models
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
Low-Rank Adaptation, LoRA, takes a different approach: it freezes the pretrained weight matrix entirely and represents the update as the product of two much smaller matrices. For a frozen weight matrix , LoRA represents the adapted weight as . Only B and A are trained; never moves during fine-tuning [ 5 ] . The effect on trainable parameter count for that one matrix is to fall from dk to r(d + k) , which is small whenever the chosen rank r is small relative to d and k — and because BA can be merged back into after training, LoRA adds no extra inference latency once deployed. Hu and colleagues report, for their comparison against full…