← All parts of this equation

Equation 6 · Part 16 · One Model, Many Modalities: What Multimodal Systems Actually Share

Numerator: exp(langle u_i, v_i rangle / τ)

L=−1B∑i=1Blog⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨ui,vj⟩/τ),\mathcal{L} = -\frac{1}{B} \sum_{i=1}^{B} \log \frac{\exp(\langle u_i, v_i \rangle / \tau)}{\sum_{j=1}^{B} \exp(\langle u_i, v_j \rangle / \tau)},
exp⁡(⟨ui,vi⟩/τ)\exp(\langle u_i, v_i \rangle / \tau)

What this part means

The complete quantity above the fraction bar.

Its job in the formula

exp(langle uiu_i, viv_i rangle / τ) occurs above the fraction bar. The numerator is divided by the entire denominator below it.

The passage around this formula

Alignment is the step that made cross-modal retrieval work, and it has an unusually clean formulation. Contrastive language–image pretraining takes a batch of B image–text pairs, encodes each side separately into a normalised vector, and trains both encoders so that matched pairs score higher than mismatched ones. In its softmax form the objective is L=−1B∑i=1Blog⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨ui,vj⟩/τ)\mathcal{L} = -\frac{1}{B} \sum_{i=1}^{B} \log \frac{\exp(\langle u_i, v_i \rangle / \tau)}{\sum_{j=1}^{B} \exp(\langle u_i, v_j \rangle / \tau)}. with image embedding uiu_i , text embedding viv_i , and a learned temperature τ\tau . Radford and colleagues showed that this simple pre-training task, applied to 400 million image–text pairs collected from the internet, matched the accuracy of the original ResNet-50 on ImageNet zero-shot without using any of the 1.28 million…

Read this part in the article →

Learn the underlying idea

A fraction a/b means a divided by b. The top number is the numerator; the bottom number is the denominator, and it cannot be zero.

Open the illustrated fractions: division written vertically guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.