Equation 7 · One Model, Many Modalities: What Multimodal Systems Actually Share
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the image embedding. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
with image embedding , text embedding , and a learned temperature . Radford and colleagues showed that this simple pre-training task, applied to 400 million image–text pairs collected from the internet, matched the accuracy of the original ResNet-50 on ImageNet zero-shot without using any of the 1.28 million training examples that network was trained on [ 1 ] . A later variant replaces the softmax with a pairwise sigmoid loss that does not require a global view of the batch’s pairwise similarities for normalisation, and the authors report training a model to 84.5% ImageNet zero-shot accuracy in two days on four TPUv4 chips, while also finding that the benefit of larger batches…
Read the full surrounding passage
with image embedding , text embedding , and a learned temperature . Radford and colleagues showed that this simple pre-training task, applied to 400 million image–text pairs collected from the internet, matched the accuracy of the original ResNet-50 on ImageNet zero-shot without using any of the 1.28 million training examples that network was trained on [ 1 ] . A later variant replaces the softmax with a pairwise sigmoid loss that does not require a global view of the batch’s pairwise similarities for normalisation, and the authors report training a model to 84.5% ImageNet zero-shot accuracy in two days on four TPUv4 chips, while also finding that the benefit of larger batches saturates around 32,000 examples [ 6 ] .
Sources cited in the surrounding passage
- [1] Learning Transferable Visual Models From Natural Language Supervision ↗
- [6] Sigmoid Loss for Language Image Pre-Training ↗
These citations give research context. Read each source to check which claims it supports.
Return to One Model, Many Modalities: What Multimodal Systems Actually Share