← All parts of this equation

Equation 6 · Part 18 · One Model, Many Modalities: What Multimodal Systems Actually Share

Starting index or lower bound: j=1

L=−1B∑i=1Blog⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨ui,vj⟩/τ),\mathcal{L} = -\frac{1}{B} \sum_{i=1}^{B} \log \frac{\exp(\langle u_i, v_i \rangle / \tau)}{\sum_{j=1}^{B} \exp(\langle u_i, v_j \rangle / \tau)},
j=1j=1

What this part means

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Its job in the formula

j=1 occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

The passage around this formula

Alignment is the step that made cross-modal retrieval work, and it has an unusually clean formulation. Contrastive language–image pretraining takes a batch of B image–text pairs, encodes each side separately into a normalised vector, and trains both encoders so that matched pairs score higher than mismatched ones. In its softmax form the objective is L=−1B∑i=1Blog⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨ui,vj⟩/τ)\mathcal{L} = -\frac{1}{B} \sum_{i=1}^{B} \log \frac{\exp(\langle u_i, v_i \rangle / \tau)}{\sum_{j=1}^{B} \exp(\langle u_i, v_j \rangle / \tau)}. with image embedding uiu_i , text embedding viv_i , and a learned temperature τ\tau . Radford and colleagues showed that this simple pre-training task, applied to 400 million image–text pairs collected from the internet, matched the accuracy of the original ResNet-50 on ImageNet zero-shot without using any of the 1.28 million…

Read this part in the article →

Learn the underlying idea

Σ adds a collection of terms. Π multiplies them. The lower and upper labels tell you which terms belong to the collection.

Open the illustrated sums and products: repeat an operation over an index guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.