Published equation contexts
Why this formula appears here
Only B and A are trained; never moves during fine-tuning [ 5 ] . The effect on trainable parameter count for that one matrix is to fall from dk to r(d + k) , which is small whenever the chosen rank r is small relative to d and k — and because BA can be merged back into after training, LoRA adds no extra inference latency once deployed. Hu and colleagues report, for their comparison against full fine-tuning of GPT-3 175B, up to a 10,000-fold reduction in trainable parameters and a threefold reduction in GPU memory requirement, with quality on par with or better than full fine-tuning on the benchmarks they tested [ 5 ] . Those figures are the paper’s own reported comparison for a…
Read the representative guide
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
Research cited beside this formula
Published contexts (1)
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 11 · Open Models
Actually Deploying an Open-Weight Model in Production
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
Only B and A are trained; never moves during fine-tuning [ 5 ] . The effect on trainable parameter count for that one matrix is to fall from dk to r(d + k) , which is small whenever the chosen rank r is small relative to d and k — and because BA can be merged back into after training, LoRA adds no extra inference latency once deployed. Hu and colleagues report, for their comparison against full fine-tuning of GPT-3 175B, up to a 10,000-fold reduction in trainable parameters and a threefold reduction in GPU memory requirement, with quality on par with or better than full fine-tuning on the benchmarks they tested [ 5 ] . Those figures are the paper’s own reported comparison for a…