Equation 6 · Part 17 · One Model, Many Modalities: What Multimodal Systems Actually Share
Denominator: sum_j=1^B exp(langle u_i, v_j rangle / τ)
What this part means
The complete quantity below the fraction bar; it must be nonzero for this division.
Its job in the formula
su=1^B exp(langle , rangle / τ) occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Full expression→Denominator: sum_j=1^B exp(langle u_i, v_j rangle / τ)→Article meaning
The passage around this formula
Alignment is the step that made cross-modal retrieval work, and it has an unusually clean formulation. Contrastive language–image pretraining takes a batch of B image–text pairs, encodes each side separately into a normalised vector, and trains both encoders so that matched pairs score higher than mismatched ones. In its softmax form the objective is . with image embedding , text embedding , and a learned temperature . Radford and colleagues showed that this simple pre-training task, applied to 400 million image–text pairs collected from the internet, matched the accuracy of the original ResNet-50 on ImageNet zero-shot without using any of the 1.28 million…
Learn the underlying idea
A fraction a/b means a divided by b. The top number is the numerator; the bottom number is the denominator, and it cannot be zero.
Open the illustrated fractions: division written vertically guide →
Sources cited in the surrounding passage
- [1] Learning Transferable Visual Models From Natural Language Supervision ↗
- [6] Sigmoid Loss for Language Image Pre-Training ↗
These citations provide research context; check each source for the exact claim it supports.