← Mathematical compendium

Published equation contexts

W=W0+ΔW=W0+BA,B∈Rd×r, A∈Rr×k, r≪min⁡(d,k)W = W_0 + \Delta W = W_0 + BA, \qquad B \in \mathbb{R}^{d \times r},\ A \in \mathbb{R}^{r \times k},\ r \ll \min(d, k)

Why this formula appears here

Low-Rank Adaptation, LoRA, takes a different approach: it freezes the pretrained weight matrix entirely and represents the update as the product of two much smaller matrices. For a frozen weight matrix W0W_0 ∈\in Rd×k\mathbb{R}^{d \times k} , LoRA represents the adapted weight as W=W0+ΔW=W0+BA,B∈Rd×r, A∈Rr×k, r≪min⁡(d,k)W = W_0 + \Delta W = W_0 + BA, \qquad B \in \mathbb{R}^{d \times r},\ A \in \mathbb{R}^{r \times k},\ r \ll \min(d, k). Only B and A are trained; W0W_0 never moves during fine-tuning [ 5 ] . The effect on trainable parameter count for that one matrix is to fall from dk to r(d + k) , which is small whenever the chosen rank r is small relative to d and k — and because BA can be merged back into W0W_0 after training, LoRA adds no extra inference latency once deployed. Hu and colleagues report, for their comparison against full…

Read the full article-specific guide →

Read the representative guide

Rd×r\mathbb{R}^{d \times r}

Symbol R^d × r

RdR^d × r appears in the objective or constraint used by the optimization on the right.

Read this term in its guide →
Rr×k\mathbb{R}^{r \times k}

Symbol R^r × k

RrR^r × k appears in the objective or constraint used by the optimization on the right.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

W=W0+ΔW=W0+BA,B∈Rd×r, A∈Rr×k, r≪min⁡(d,k)W = W_0 + \Delta W = W_0 + BA, \qquad B \in \mathbb{R}^{d \times r},\ A \in \mathbb{R}^{r \times k},\ r \ll \min(d, k)

Equation 2 · Open Models

Actually Deploying an Open-Weight Model in Production

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Low-Rank Adaptation, LoRA, takes a different approach: it freezes the pretrained weight matrix entirely and represents the update as the product of two much smaller matrices. For a frozen weight matrix W0W_0 ∈\in Rd×k\mathbb{R}^{d \times k} , LoRA represents the adapted weight as W=W0+ΔW=W0+BA,B∈Rd×r, A∈Rr×k, r≪min⁡(d,k)W = W_0 + \Delta W = W_0 + BA, \qquad B \in \mathbb{R}^{d \times r},\ A \in \mathbb{R}^{r \times k},\ r \ll \min(d, k). Only B and A are trained; W0W_0 never moves during fine-tuning [ 5 ] . The effect on trainable parameter count for that one matrix is to fall from dk to r(d + k) , which is small whenever the chosen rank r is small relative to d and k — and because BA can be merged back into W0W_0 after training, LoRA adds no extra inference latency once deployed. Hu and colleagues report, for their comparison against full…

Meanings in this article

Equation guide → · Article →