← Back to article

Equation 4 · Actually Deploying an Open-Weight Model in Production

What does this equation mean?

AA

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

trained. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

AA

Symbol A

trained.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Only B and A are trained; W0W_0 never moves during fine-tuning [ 5 ] . The effect on trainable parameter count for that one matrix is to fall from dk to r(d + k) , which is small whenever the chosen rank r is small relative to d and k — and because BA can be merged back into W0W_0 after training, LoRA adds no extra inference latency once deployed. Hu and colleagues report, for their comparison against full fine-tuning of GPT-3 175B, up to a 10,000-fold reduction in trainable parameters and a threefold reduction in GPU memory requirement, with quality on par with or better than full fine-tuning on the benchmarks they tested [ 5 ] . Those figures are the paper’s own reported comparison for a…
Read the full surrounding passage
Only B and A are trained; W0W_0 never moves during fine-tuning [ 5 ] . The effect on trainable parameter count for that one matrix is to fall from dk to r(d + k) , which is small whenever the chosen rank r is small relative to d and k — and because BA can be merged back into W0W_0 after training, LoRA adds no extra inference latency once deployed. Hu and colleagues report, for their comparison against full fine-tuning of GPT-3 175B, up to a 10,000-fold reduction in trainable parameters and a threefold reduction in GPU memory requirement, with quality on par with or better than full fine-tuning on the benchmarks they tested [ 5 ] . Those figures are the paper’s own reported comparison for a specific model and are not a guaranteed ratio for every architecture or rank choice, but the underlying mechanism — freezing the base and training a small low-rank update — is what makes LoRA the practical default for narrower behavioral adjustments: the adapter is small enough to store, version, and swap independently of the frozen base, and several task-specific adapters can share one base checkpoint in production rather than each requiring a full duplicate of the model.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Actually Deploying an Open-Weight Model in Production

Browse the mathematical compendium →