← Mathematical compendium

Published equation contexts

r(d+k)r(d + k)

Why this formula appears here

Only B and A are trained; W0W_0 never moves during fine-tuning [ 5 ] . The effect on trainable parameter count for that one matrix is to fall from dk to r(d + k) , which is small whenever the chosen rank r is small relative to d and k — and because BA can be merged back into W0W_0 after training, LoRA adds no extra inference latency once deployed. Hu and colleagues report, for their comparison against full fine-tuning of GPT-3 175B, up to a 10,000-fold reduction in trainable parameters and a threefold reduction in GPU memory requirement, with quality on par with or better than full fine-tuning on the benchmarks they tested [ 5 ] . Those figures are the paper’s own reported comparison for a…

Read the full article-specific guide →

Read the representative guide

rr

Symbol r

r is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
dd

Symbol d

d is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
kk

Symbol k

k is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

r(d+k)r(d + k)

Equation 7 · Open Models

Actually Deploying an Open-Weight Model in Production

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

Only B and A are trained; W0W_0 never moves during fine-tuning [ 5 ] . The effect on trainable parameter count for that one matrix is to fall from dk to r(d + k) , which is small whenever the chosen rank r is small relative to d and k — and because BA can be merged back into W0W_0 after training, LoRA adds no extra inference latency once deployed. Hu and colleagues report, for their comparison against full fine-tuning of GPT-3 175B, up to a 10,000-fold reduction in trainable parameters and a threefold reduction in GPU memory requirement, with quality on par with or better than full fine-tuning on the benchmarks they tested [ 5 ] . Those figures are the paper’s own reported comparison for a…

Equation guide → · Article →