← All parts of this equation

Equation 5 · Part 12 · A History of Multimodal AI

addition

L=−12B∑i=1B[log⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨ui,vj⟩/τ)+log⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨uj,vi⟩/τ)],\mathcal{L} = -\frac{1}{2B}\sum_{i=1}^{B}\left[\log\frac{\exp(\langle u_i, v_i\rangle/\tau)}{\sum_{j=1}^{B}\exp(\langle u_i, v_j\rangle/\tau)} + \log\frac{\exp(\langle u_i, v_i\rangle/\tau)}{\sum_{j=1}^{B}\exp(\langle u_j, v_i\rangle/\tau)}\right],
addition

What this part means

Add the term after the plus sign to the term or group before it.

Its job in the formula

Add the term after the plus sign to the term or group before it.

The passage around this formula

Radford and colleagues’ 2021 paper, “Learning Transferable Visual Models From Natural Language Supervision,” is usually remembered simply as CLIP, and the compression loses the specific thing that made it a turning point rather than an incremental improvement. The method itself is not exotic: encode an image, encode its paired caption, and train both encoders so that the true pairing scores higher than every mismatched pairing drawn from the same batch. For a batch of B image-text pairs with normalised image embedding uiu_i , text embedding viv_i , and a learned temperature τ\tau , the symmetric form of the objective is L=−12B∑i=1B[log⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨ui,vj⟩/τ)+log⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨uj,vi⟩/τ)]\mathcal{L} = -\frac{1}{2B}\sum_{i=1}^{B}\left[\log\frac{\exp(\langle u_i, v_i\rangle/\tau)}{\sum_{j=1}^{B}\exp(\langle u_i, v_j\rangle/\tau)} + \log\frac{\exp(\langle u_i, v_i\rangle/\tau)}{\sum_{j=1}^{B}\exp(\langle u_j, v_i\rangle/\tau)}\right]. an image-to-text and a text-to-image cross-entropy…

Read this part in the article →

Learn the underlying idea

Addition combines quantities; subtraction measures the signed difference between them. Parentheses show what is combined before the rest of the expression is evaluated.

Open the illustrated addition and subtraction in an equation guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.