← Back to article

Equation 2 · Actually Deploying an Open-Weight Model in Production

What does this equation mean?

W=W0+ΔW=W0+BA,B∈Rd×r, A∈Rr×k, r≪min⁡(d,k)W = W_0 + \Delta W = W_0 + BA, \qquad B \in \mathbb{R}^{d \times r},\ A \in \mathbb{R}^{r \times k},\ r \ll \min(d, k)

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

WW

Symbol W

W is the quantity selected or evaluated by the optimization written on the right.

Understand this part →

W0W_0

Symbol W_0

W0W_0 appears in the objective or constraint used by the optimization on the right.

Understand this part →

ΔW\Delta W

Symbol Δ W

Δ W appears in the objective or constraint used by the optimization on the right.

Understand this part →

BB

Symbol B

the only.

Understand this part →

AA

Symbol A

trained.

Understand this part →

Rd×r\mathbb{R}^{d \times r}

Symbol R^d × r

RdR^d × r appears in the objective or constraint used by the optimization on the right.

Understand this part →

Rr×k\mathbb{R}^{r \times k}

Symbol R^r × k

RrR^r × k appears in the objective or constraint used by the optimization on the right.

Understand this part →

rr

Symbol r

r appears in the objective or constraint used by the optimization on the right.

Understand this part →

dd

Symbol d

d appears in the objective or constraint used by the optimization on the right.

Understand this part →

kk

Symbol k

k appears in the objective or constraint used by the optimization on the right.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
multiplication

multiplication

Multiply the quantities on either side.

Understand this part →

addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

change

change

Capital delta attached to a quantity marks a difference between two values of that quantity; the article’s sign convention determines the order.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Low-Rank Adaptation, LoRA, takes a different approach: it freezes the pretrained weight matrix entirely and represents the update as the product of two much smaller matrices. For a frozen weight matrix W0W_0 ∈\in Rd×k\mathbb{R}^{d \times k} , LoRA represents the adapted weight as W=W0+ΔW=W0+BA,B∈Rd×r, A∈Rr×k, r≪min⁡(d,k)W = W_0 + \Delta W = W_0 + BA, \qquad B \in \mathbb{R}^{d \times r},\ A \in \mathbb{R}^{r \times k},\ r \ll \min(d, k). Only B and A are trained; W0W_0 never moves during fine-tuning [ 5 ] . The effect on trainable parameter count for that one matrix is to fall from dk to r(d + k) , which is small whenever the chosen rank r is small relative to d and k — and because BA can be merged back into W0W_0 after training, LoRA adds no extra inference latency once deployed. Hu and colleagues report, for their comparison against full…
Read the full surrounding passage
Low-Rank Adaptation, LoRA, takes a different approach: it freezes the pretrained weight matrix entirely and represents the update as the product of two much smaller matrices. For a frozen weight matrix W0W_0 ∈\in Rd×k\mathbb{R}^{d \times k} , LoRA represents the adapted weight as W=W0+ΔW=W0+BA,B∈Rd×r, A∈Rr×k, r≪min⁡(d,k)W = W_0 + \Delta W = W_0 + BA, \qquad B \in \mathbb{R}^{d \times r},\ A \in \mathbb{R}^{r \times k},\ r \ll \min(d, k). Only B and A are trained; W0W_0 never moves during fine-tuning [ 5 ] . The effect on trainable parameter count for that one matrix is to fall from dk to r(d + k) , which is small whenever the chosen rank r is small relative to d and k — and because BA can be merged back into W0W_0 after training, LoRA adds no extra inference latency once deployed. Hu and colleagues report, for their comparison against full fine-tuning of GPT-3 175B, up to a 10,000-fold reduction in trainable parameters and a threefold reduction in GPU memory requirement, with quality on par with or better than full fine-tuning on the benchmarks they tested [ 5 ] . Those figures are the paper’s own reported comparison for a specific model and are not a guaranteed ratio for every architecture or rank choice, but the underlying mechanism — freezing the base and training a small low-rank update — is what makes LoRA the practical default for narrower behavioral adjustments: the adapter is small enough to store, version, and swap independently of the frozen base, and several task-specific adapters can share one base checkpoint in production rather than each requiring a full duplicate of the model.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Actually Deploying an Open-Weight Model in Production

See this formula across 1 published context →

Browse the mathematical compendium →