← Mathematical compendium

Published equation contexts

L=−1B∑i=1Blog⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨ui,vj⟩/τ)\mathcal{L} = -\frac{1}{B} \sum_{i=1}^{B} \log \frac{\exp(\langle u_i, v_i \rangle / \tau)}{\sum_{j=1}^{B} \exp(\langle u_i, v_j \rangle / \tau)}

Why this formula appears here

Alignment is the step that made cross-modal retrieval work, and it has an unusually clean formulation. Contrastive language–image pretraining takes a batch of B image–text pairs, encodes each side separately into a normalised vector, and trains both encoders so that matched pairs score higher than mismatched ones. In its softmax form the objective is L=−1B∑i=1Blog⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨ui,vj⟩/τ)\mathcal{L} = -\frac{1}{B} \sum_{i=1}^{B} \log \frac{\exp(\langle u_i, v_i \rangle / \tau)}{\sum_{j=1}^{B} \exp(\langle u_i, v_j \rangle / \tau)}. with image embedding uiu_i , text embedding viv_i , and a learned temperature τ\tau . Radford and colleagues showed that this simple pre-training task, applied to 400 million image–text pairs collected from the internet, matched the accuracy of the original ResNet-50 on ImageNet zero-shot without using any of the 1.28 million…

Read the full article-specific guide →

Read the representative guide

BB

Symbol B

B occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →
ii

Symbol i

i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Read this term in its guide →
jj

Symbol j

j occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →
vjv_j

Symbol v_j

vjv_j occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →
i=1i=1

Starting index or lower bound: i=1

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Read this term in its guide →
BB

Ending index or upper bound: B

This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.

Read this term in its guide →
exp⁡(⟨ui,vi⟩/τ)\exp(\langle u_i, v_i \rangle / \tau)

Numerator: exp(langle u_i, v_i rangle / τ)

The complete quantity above the fraction bar.

Read this term in its guide →
∑j=1Bexp⁡(⟨ui,vj⟩/τ)\sum_{j=1}^{B} \exp(\langle u_i, v_j \rangle / \tau)

Denominator: sum_j=1^B exp(langle u_i, v_j rangle / τ)

The complete quantity below the fraction bar; it must be nonzero for this division.

Read this term in its guide →
j=1j=1

Starting index or lower bound: j=1

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Read this term in its guide →
BB

Ending index or upper bound: B

This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

L=−1B∑i=1Blog⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨ui,vj⟩/τ),\mathcal{L} = -\frac{1}{B} \sum_{i=1}^{B} \log \frac{\exp(\langle u_i, v_i \rangle / \tau)}{\sum_{j=1}^{B} \exp(\langle u_i, v_j \rangle / \tau)},

Equation 6 · Foundation Models

One Model, Many Modalities: What Multimodal Systems Actually Share

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Alignment is the step that made cross-modal retrieval work, and it has an unusually clean formulation. Contrastive language–image pretraining takes a batch of B image–text pairs, encodes each side separately into a normalised vector, and trains both encoders so that matched pairs score higher than mismatched ones. In its softmax form the objective is L=−1B∑i=1Blog⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨ui,vj⟩/τ)\mathcal{L} = -\frac{1}{B} \sum_{i=1}^{B} \log \frac{\exp(\langle u_i, v_i \rangle / \tau)}{\sum_{j=1}^{B} \exp(\langle u_i, v_j \rangle / \tau)}. with image embedding uiu_i , text embedding viv_i , and a learned temperature τ\tau . Radford and colleagues showed that this simple pre-training task, applied to 400 million image–text pairs collected from the internet, matched the accuracy of the original ResNet-50 on ImageNet zero-shot without using any of the 1.28 million…

Meanings in this article

Equation guide → · Article →