← All parts of this equation

Equation 3 · Part 1 · A History of Multimodal AI

Symbol v_i

viv_i
viv_i

What this part means

the text embedding.

Its job in the formula

viv_i is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Where the article explains it

For a batch of B image-text pairs with normalised image embedding uiu_i , text embedding viv_i , and a learned temperature τ\tau , the symmetric form of the objective is

The passage around this formula

Radford and colleagues’ 2021 paper, “Learning Transferable Visual Models From Natural Language Supervision,” is usually remembered simply as CLIP, and the compression loses the specific thing that made it a turning point rather than an incremental improvement. The method itself is not exotic: encode an image, encode its paired caption, and train both encoders so that the true pairing scores higher than every mismatched pairing drawn from the same batch. For a batch of B image-text pairs with normalised image embedding uiu_i , text embedding viv_i , and a learned temperature τ\tau , the symmetric form of the objective is

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.