Published equation contexts
Why this formula appears here
Alignment is the step that made cross-modal retrieval work, and it has an unusually clean formulation. Contrastive language–image pretraining takes a batch of B image–text pairs, encodes each side separately into a normalised vector, and trains both encoders so that matched pairs score higher than mismatched ones. In its softmax form the objective is . with image embedding , text embedding , and a learned temperature . Radford and colleagues showed that this simple pre-training task, applied to 400 million image–text pairs collected from the internet, matched the accuracy of the original ResNet-50 on ImageNet zero-shot without using any of the 1.28 million…
Read the representative guide
Symbol B
B occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Read this term in its guide →Symbol i
i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Read this term in its guide →Symbol j
j occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Read this term in its guide →Symbol v_j
occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Read this term in its guide →Starting index or lower bound: i=1
This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
Read this term in its guide →Ending index or upper bound: B
This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.
Read this term in its guide →Numerator: exp(langle u_i, v_i rangle / τ)
The complete quantity above the fraction bar.
Read this term in its guide →Denominator: sum_j=1^B exp(langle u_i, v_j rangle / τ)
The complete quantity below the fraction bar; it must be nonzero for this division.
Read this term in its guide →Starting index or lower bound: j=1
This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
Read this term in its guide →Ending index or upper bound: B
This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.
Read this term in its guide →How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
Research cited beside this formula
Published contexts (1)
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 6 · Foundation Models
One Model, Many Modalities: What Multimodal Systems Actually Share
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
Alignment is the step that made cross-modal retrieval work, and it has an unusually clean formulation. Contrastive language–image pretraining takes a batch of B image–text pairs, encodes each side separately into a normalised vector, and trains both encoders so that matched pairs score higher than mismatched ones. In its softmax form the objective is . with image embedding , text embedding , and a learned temperature . Radford and colleagues showed that this simple pre-training task, applied to 400 million image–text pairs collected from the internet, matched the accuracy of the original ResNet-50 on ImageNet zero-shot without using any of the 1.28 million…
Meanings in this article
- : In its softmax form the objective.
- : the image embedding.
- : the text embedding.
- : the learned temperature.