← Mathematical compendium

Published equation contexts

ΔW=BA,B∈Rd×r, A∈Rr×k, r≪min⁡(d,k)\Delta W = BA, \qquad B \in \mathbb{R}^{d \times r},\ A \in \mathbb{R}^{r \times k},\ r \ll \min(d, k)

Why this formula appears here

When the adapter route is chosen, the engineering default within it is equally clear: adapt, do not retrain. Low-Rank Adaptation freezes the pretrained weights and injects a pair of small trainable matrices into selected layers, so that a weight update is expressed as a low-rank product rather than a dense matrix the size of the original layer, ΔW=BA,B∈Rd×r, A∈Rr×k, r≪min⁡(d,k)\Delta W = BA, \qquad B \in \mathbb{R}^{d \times r},\ A \in \mathbb{R}^{r \times k},\ r \ll \min(d, k). with the forward pass computing h = W0W_0 x + Δ\Delta W x against the frozen base weight W0W_0 . The method’s authors report reducing the number of trainable parameters by roughly ten thousand times and GPU memory requirements by roughly three times relative to full fine-tuning of a 175-billion-parameter model, while matching or…

Read the full article-specific guide →

Read the representative guide

Rd×r\mathbb{R}^{d \times r}

Symbol R^d × r

RdR^d × r appears in the objective or constraint used by the optimization on the right.

Read this term in its guide →
Rr×k\mathbb{R}^{r \times k}

Symbol R^r × k

RrR^r × k appears in the objective or constraint used by the optimization on the right.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

ΔW=BA,B∈Rd×r, A∈Rr×k, r≪min⁡(d,k),\Delta W = BA, \qquad B \in \mathbb{R}^{d \times r},\ A \in \mathbb{R}^{r \times k},\ r \ll \min(d, k),

Equation 1 · Foundation Models

Building a Multimodal AI Application That Actually Uses Its Inputs

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

When the adapter route is chosen, the engineering default within it is equally clear: adapt, do not retrain. Low-Rank Adaptation freezes the pretrained weights and injects a pair of small trainable matrices into selected layers, so that a weight update is expressed as a low-rank product rather than a dense matrix the size of the original layer, ΔW=BA,B∈Rd×r, A∈Rr×k, r≪min⁡(d,k)\Delta W = BA, \qquad B \in \mathbb{R}^{d \times r},\ A \in \mathbb{R}^{r \times k},\ r \ll \min(d, k). with the forward pass computing h = W0W_0 x + Δ\Delta W x against the frozen base weight W0W_0 . The method’s authors report reducing the number of trainable parameters by roughly ten thousand times and GPU memory requirements by roughly three times relative to full fine-tuning of a 175-billion-parameter model, while matching or…

Equation guide → · Article →