Equation 6 · Part 9 · One Model, Many Modalities: What Multimodal Systems Actually Share
=
=
What this part means
The expressions on both sides represent the same quantity under the stated assumptions.
Its job in the formula
The equals sign connects the complete expression on the left with the complete expression on the right. Both sides must have compatible units.
Full expression→=→Article meaning
The passage around this formula
Alignment is the step that made cross-modal retrieval work, and it has an unusually clean formulation. Contrastive language–image pretraining takes a batch of B image–text pairs, encodes each side separately into a normalised vector, and trains both encoders so that matched pairs score higher than mismatched ones. In its softmax form the objective is . with image embedding , text embedding , and a learned temperature . Radford and colleagues showed that this simple pre-training task, applied to 400 million image–text pairs collected from the internet, matched the accuracy of the original ResNet-50 on ImageNet zero-shot without using any of the 1.28 million…
Learn the underlying idea
An equals sign says that the expression on its left and the expression on its right have the same value under the stated definitions and assumptions.
Open the illustrated equality: what the equals sign claims guide →
Sources cited in the surrounding passage
- [1] Learning Transferable Visual Models From Natural Language Supervision ↗
- [6] Sigmoid Loss for Language Image Pre-Training ↗
These citations provide research context; check each source for the exact claim it supports.