Equation 5 · Part 1 · A History of Multimodal AI
Symbol L
What this part means
the symmetric form of the objective.
Its job in the formula
L is part of the quantity the equation computes from the expression on the right.
Full expression→Symbol L→Article meaning
Where the article explains it
For a batch of B image-text pairs with normalised image embedding , text embedding , and a learned temperature , the symmetric form of the objective is
The passage around this formula
Radford and colleagues’ 2021 paper, “Learning Transferable Visual Models From Natural Language Supervision,” is usually remembered simply as CLIP, and the compression loses the specific thing that made it a turning point rather than an incremental improvement. The method itself is not exotic: encode an image, encode its paired caption, and train both encoders so that the true pairing scores higher than every mismatched pairing drawn from the same batch. For a batch of B image-text pairs with normalised image embedding , text embedding , and a learned temperature , the symmetric form of the objective is . an image-to-text and a text-to-image cross-entropy…
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.