← Back to article

Equation 1 · Building a Multimodal AI Application That Actually Uses Its Inputs

What does this equation mean?

ΔW=BA,B∈Rd×r, A∈Rr×k, r≪min⁡(d,k),\Delta W = BA, \qquad B \in \mathbb{R}^{d \times r},\ A \in \mathbb{R}^{r \times k},\ r \ll \min(d, k),

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationsBA, qquad B in R^d × r, A in R^r × k, r ll min(d, k)
Result or conditionΔ W
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

ΔW\Delta W

Symbol Δ W

Δ W is the quantity selected or evaluated by the optimization written on the right.

Understand this part →

BB

Symbol B

B appears in the objective or constraint used by the optimization on the right.

Understand this part →

AA

Symbol A

A appears in the objective or constraint used by the optimization on the right.

Understand this part →

Rd×r\mathbb{R}^{d \times r}

Symbol R^d × r

RdR^d × r appears in the objective or constraint used by the optimization on the right.

Understand this part →

Rr×k\mathbb{R}^{r \times k}

Symbol R^r × k

RrR^r × k appears in the objective or constraint used by the optimization on the right.

Understand this part →

rr

Symbol r

r appears in the objective or constraint used by the optimization on the right.

Understand this part →

dd

Symbol d

d appears in the objective or constraint used by the optimization on the right.

Understand this part →

kk

Symbol k

k appears in the objective or constraint used by the optimization on the right.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
multiplication

multiplication

Multiply the quantities on either side.

Understand this part →

change

change

Capital delta attached to a quantity marks a difference between two values of that quantity; the article’s sign convention determines the order.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

When the adapter route is chosen, the engineering default within it is equally clear: adapt, do not retrain. Low-Rank Adaptation freezes the pretrained weights and injects a pair of small trainable matrices into selected layers, so that a weight update is expressed as a low-rank product rather than a dense matrix the size of the original layer, ΔW=BA,B∈Rd×r, A∈Rr×k, r≪min⁡(d,k)\Delta W = BA, \qquad B \in \mathbb{R}^{d \times r},\ A \in \mathbb{R}^{r \times k},\ r \ll \min(d, k). with the forward pass computing h = W0W_0 x + Δ\Delta W x against the frozen base weight W0W_0 . The method’s authors report reducing the number of trainable parameters by roughly ten thousand times and GPU memory requirements by roughly three times relative to full fine-tuning of a 175-billion-parameter model, while matching or…
Read the full surrounding passage
When the adapter route is chosen, the engineering default within it is equally clear: adapt, do not retrain. Low-Rank Adaptation freezes the pretrained weights and injects a pair of small trainable matrices into selected layers, so that a weight update is expressed as a low-rank product rather than a dense matrix the size of the original layer, ΔW=BA,B∈Rd×r, A∈Rr×k, r≪min⁡(d,k)\Delta W = BA, \qquad B \in \mathbb{R}^{d \times r},\ A \in \mathbb{R}^{r \times k},\ r \ll \min(d, k). with the forward pass computing h = W0W_0 x + Δ\Delta W x against the frozen base weight W0W_0 . The method’s authors report reducing the number of trainable parameters by roughly ten thousand times and GPU memory requirements by roughly three times relative to full fine-tuning of a 175-billion-parameter model, while matching or exceeding full fine-tuning quality and adding no additional inference latency, since the low-rank update can be merged back into the base weight at deployment time [ 1 ] . QLoRA extends the same idea to quantized base weights, and its authors report finetuning a 65-billion-parameter model on a single 48-gigabyte GPU while preserving full 16-bit finetuning performance, with their best resulting model reaching 99.3 percent of a reference chat model’s quality after twenty-four hours of finetuning on one GPU [ 2 ] . Hugging Face’s PEFT library packages LoRA and related adapter methods as a standard toolchain specifically so that adapting a large pretrained model no longer requires updating all of its parameters, which the library’s own documentation describes as “prohibitively costly” for most teams, while integrating directly with the standard training and inference stack [ 10 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Building a Multimodal AI Application That Actually Uses Its Inputs

See this formula across 1 published context →

Browse the mathematical compendium →