← Back to article

Equation 7 · One Model, Many Modalities: What Multimodal Systems Actually Share

What does this equation mean?

uiu_i

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

the image embedding. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

uiu_i

Symbol u_i

the image embedding.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

with image embedding uiu_i , text embedding viv_i , and a learned temperature τ\tau . Radford and colleagues showed that this simple pre-training task, applied to 400 million image–text pairs collected from the internet, matched the accuracy of the original ResNet-50 on ImageNet zero-shot without using any of the 1.28 million training examples that network was trained on [ 1 ] . A later variant replaces the softmax with a pairwise sigmoid loss that does not require a global view of the batch’s pairwise similarities for normalisation, and the authors report training a model to 84.5% ImageNet zero-shot accuracy in two days on four TPUv4 chips, while also finding that the benefit of larger batches…
Read the full surrounding passage
with image embedding uiu_i , text embedding viv_i , and a learned temperature τ\tau . Radford and colleagues showed that this simple pre-training task, applied to 400 million image–text pairs collected from the internet, matched the accuracy of the original ResNet-50 on ImageNet zero-shot without using any of the 1.28 million training examples that network was trained on [ 1 ] . A later variant replaces the softmax with a pairwise sigmoid loss that does not require a global view of the batch’s pairwise similarities for normalisation, and the authors report training a model to 84.5% ImageNet zero-shot accuracy in two days on four TPUv4 chips, while also finding that the benefit of larger batches saturates around 32,000 examples [ 6 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to One Model, Many Modalities: What Multimodal Systems Actually Share

Browse the mathematical compendium →