← All parts of this equation

Equation 6 · Part 3 · One Model, Many Modalities: What Multimodal Systems Actually Share

Symbol i

L=−1B∑i=1Blog⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨ui,vj⟩/τ),\mathcal{L} = -\frac{1}{B} \sum_{i=1}^{B} \log \frac{\exp(\langle u_i, v_i \rangle / \tau)}{\sum_{j=1}^{B} \exp(\langle u_i, v_j \rangle / \tau)},
ii

What this part means

i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Its job in the formula

i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

The passage around this formula

Alignment is the step that made cross-modal retrieval work, and it has an unusually clean formulation. Contrastive language–image pretraining takes a batch of B image–text pairs, encodes each side separately into a normalised vector, and trains both encoders so that matched pairs score higher than mismatched ones. In its softmax form the objective is L=−1B∑i=1Blog⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨ui,vj⟩/τ)\mathcal{L} = -\frac{1}{B} \sum_{i=1}^{B} \log \frac{\exp(\langle u_i, v_i \rangle / \tau)}{\sum_{j=1}^{B} \exp(\langle u_i, v_j \rangle / \tau)}. with image embedding uiu_i , text embedding viv_i , and a learned temperature τ\tau . Radford and colleagues showed that this simple pre-training task, applied to 400 million image–text pairs collected from the internet, matched the accuracy of the original ResNet-50 on ImageNet zero-shot without using any of the 1.28 million…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.