Equation 6 · Part 8 · One Model, Many Modalities: What Multimodal Systems Actually Share
Symbol v_j
What this part means
occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Its job in the formula
occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Full expression→Symbol v_j→Article meaning
The passage around this formula
Alignment is the step that made cross-modal retrieval work, and it has an unusually clean formulation. Contrastive language–image pretraining takes a batch of B image–text pairs, encodes each side separately into a normalised vector, and trains both encoders so that matched pairs score higher than mismatched ones. In its softmax form the objective is . with image embedding , text embedding , and a learned temperature . Radford and colleagues showed that this simple pre-training task, applied to 400 million image–text pairs collected from the internet, matched the accuracy of the original ResNet-50 on ImageNet zero-shot without using any of the 1.28 million…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the surrounding passage
- [1] Learning Transferable Visual Models From Natural Language Supervision ↗
- [6] Sigmoid Loss for Language Image Pre-Training ↗
These citations provide research context; check each source for the exact claim it supports.