Equation 3 · Part 1 · A History of Multimodal AI
Symbol v_i
What this part means
the text embedding.
Its job in the formula
is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol v_i→Article meaning
Where the article explains it
For a batch of B image-text pairs with normalised image embedding , text embedding , and a learned temperature , the symmetric form of the objective is
The passage around this formula
Radford and colleagues’ 2021 paper, “Learning Transferable Visual Models From Natural Language Supervision,” is usually remembered simply as CLIP, and the compression loses the specific thing that made it a turning point rather than an incremental improvement. The method itself is not exotic: encode an image, encode its paired caption, and train both encoders so that the true pairing scores higher than every mismatched pairing drawn from the same batch. For a batch of B image-text pairs with normalised image embedding , text embedding , and a learned temperature , the symmetric form of the objective is
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the article section
These citations provide research context; check each source for the exact claim it supports.