← Back to article

Equation 5 · A History of Multimodal AI

What does this equation mean?

L=−12B∑i=1B[log⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨ui,vj⟩/τ)+log⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨uj,vi⟩/τ)],\mathcal{L} = -\frac{1}{2B}\sum_{i=1}^{B}\left[\log\frac{\exp(\langle u_i, v_i\rangle/\tau)}{\sum_{j=1}^{B}\exp(\langle u_i, v_j\rangle/\tau)} + \log\frac{\exp(\langle u_i, v_i\rangle/\tau)}{\sum_{j=1}^{B}\exp(\langle u_j, v_i\rangle/\tau)}\right],

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Start with1
Divide by2B
This relates toL
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

L\mathcal{L}

Symbol L

the symmetric form of the objective.

Understand this part →

BB

Symbol B

B occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Understand this part →

ii

Symbol i

i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Understand this part →

uiu_i

Symbol u_i

the normalised image embedding.

Understand this part →

viv_i

Symbol v_i

the text embedding.

Understand this part →

τ\tau

Symbol τ

the learned temperature.

Understand this part →

jj

Symbol j

j occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Understand this part →

vjv_j

Symbol v_j

vjv_j occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Understand this part →

uju_j

Symbol u_j

uju_j occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →
11

Numerator: 1

The complete quantity above the fraction bar.

Understand this part →

2B2B

Denominator: 2B

The complete quantity below the fraction bar; it must be nonzero for this division.

Understand this part →

i=1i=1

Starting index or lower bound: i=1

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Understand this part →

BB

Ending index or upper bound: B

This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.

Understand this part →

exp⁡(⟨ui,vi⟩/τ)\exp(\langle u_i, v_i\rangle/\tau)

Numerator: exp(langle u_i, v_irangle/τ)

The complete quantity above the fraction bar.

Understand this part →

See an illustrated explanation →
∑j=1Bexp⁡(⟨ui,vj⟩/τ)\sum_{j=1}^{B}\exp(\langle u_i, v_j\rangle/\tau)

Denominator: sum_j=1^Bexp(langle u_i, v_jrangle/τ)

The complete quantity below the fraction bar; it must be nonzero for this division.

Understand this part →

See an illustrated explanation →
j=1j=1

Starting index or lower bound: j=1

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Understand this part →

BB

Ending index or upper bound: B

This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.

Understand this part →

exp⁡(⟨ui,vi⟩/τ)\exp(\langle u_i, v_i\rangle/\tau)

Numerator: exp(langle u_i, v_irangle/τ)

The complete quantity above the fraction bar.

Understand this part →

See an illustrated explanation →
∑j=1Bexp⁡(⟨uj,vi⟩/τ)\sum_{j=1}^{B}\exp(\langle u_j, v_i\rangle/\tau)

Denominator: sum_j=1^Bexp(langle u_j, v_irangle/τ)

The complete quantity below the fraction bar; it must be nonzero for this division.

Understand this part →

See an illustrated explanation →
j=1j=1

Starting index or lower bound: j=1

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Understand this part →

BB

Ending index or upper bound: B

This label says where the repeated addition, multiplication, or accumulation stops. It sets the last term or end of the range.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Radford and colleagues’ 2021 paper, “Learning Transferable Visual Models From Natural Language Supervision,” is usually remembered simply as CLIP, and the compression loses the specific thing that made it a turning point rather than an incremental improvement. The method itself is not exotic: encode an image, encode its paired caption, and train both encoders so that the true pairing scores higher than every mismatched pairing drawn from the same batch. For a batch of B image-text pairs with normalised image embedding uiu_i , text embedding viv_i , and a learned temperature τ\tau , the symmetric form of the objective is L=−12B∑i=1B[log⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨ui,vj⟩/τ)+log⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨uj,vi⟩/τ)]\mathcal{L} = -\frac{1}{2B}\sum_{i=1}^{B}\left[\log\frac{\exp(\langle u_i, v_i\rangle/\tau)}{\sum_{j=1}^{B}\exp(\langle u_i, v_j\rangle/\tau)} + \log\frac{\exp(\langle u_i, v_i\rangle/\tau)}{\sum_{j=1}^{B}\exp(\langle u_j, v_i\rangle/\tau)}\right]. an image-to-text and a text-to-image cross-entropy…
Read the full surrounding passage
Radford and colleagues’ 2021 paper, “Learning Transferable Visual Models From Natural Language Supervision,” is usually remembered simply as CLIP, and the compression loses the specific thing that made it a turning point rather than an incremental improvement. The method itself is not exotic: encode an image, encode its paired caption, and train both encoders so that the true pairing scores higher than every mismatched pairing drawn from the same batch. For a batch of B image-text pairs with normalised image embedding uiu_i , text embedding viv_i , and a learned temperature τ\tau , the symmetric form of the objective is L=−12B∑i=1B[log⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨ui,vj⟩/τ)+log⁡exp⁡(⟨ui,vi⟩/τ)∑j=1Bexp⁡(⟨uj,vi⟩/τ)]\mathcal{L} = -\frac{1}{2B}\sum_{i=1}^{B}\left[\log\frac{\exp(\langle u_i, v_i\rangle/\tau)}{\sum_{j=1}^{B}\exp(\langle u_i, v_j\rangle/\tau)} + \log\frac{\exp(\langle u_i, v_i\rangle/\tau)}{\sum_{j=1}^{B}\exp(\langle u_j, v_i\rangle/\tau)}\right]. an image-to-text and a text-to-image cross-entropy averaged together, each treating every other pairing in the batch as a negative. Nothing in that objective is architecturally new; matching objectives had been used for retrieval before CLIP. What changed was scale, and what the scale was spent on: 400 million image-text pairs collected from the public internet, with no hand-assigned category label anywhere in the pipeline [ 5 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to A History of Multimodal AI

See this formula across 1 published context →

Browse the mathematical compendium →