Equation 5 · Part 1 · One Model, Many Modalities: What Multimodal Systems Actually Share
Symbol B
What this part means
B is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Its job in the formula
B is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol B→Article meaning
The passage around this formula
Alignment is the step that made cross-modal retrieval work, and it has an unusually clean formulation. Contrastive language–image pretraining takes a batch of B image–text pairs, encodes each side separately into a normalised vector, and trains both encoders so that matched pairs score higher than mismatched ones. In its softmax form the objective is
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the article section
- [1] Learning Transferable Visual Models From Natural Language Supervision ↗
- [6] Sigmoid Loss for Language Image Pre-Training ↗
- [15] When and why vision-language models behave like bags-of-words, and what to do about it? ↗
- [14] Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality ↗
These citations provide research context; check each source for the exact claim it supports.