Equation 6 · Part 3 · One Model, Many Modalities: What Multimodal Systems Actually Share
Symbol i
What this part means
i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Its job in the formula
i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Full expression→Symbol i→Article meaning
The passage around this formula
Alignment is the step that made cross-modal retrieval work, and it has an unusually clean formulation. Contrastive language–image pretraining takes a batch of B image–text pairs, encodes each side separately into a normalised vector, and trains both encoders so that matched pairs score higher than mismatched ones. In its softmax form the objective is . with image embedding , text embedding , and a learned temperature . Radford and colleagues showed that this simple pre-training task, applied to 400 million image–text pairs collected from the internet, matched the accuracy of the original ResNet-50 on ImageNet zero-shot without using any of the 1.28 million…
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the surrounding passage
- [1] Learning Transferable Visual Models From Natural Language Supervision ↗
- [6] Sigmoid Loss for Language Image Pre-Training ↗
These citations provide research context; check each source for the exact claim it supports.