← Back to article

Equation 6 · Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared

What does this equation mean?

Ljoint(θ)=E(xtext, ximg, xaud)∼D[ Ltext(θ)+Limg(θ)+Laud(θ) ]\mathcal{L}_{\text{joint}}(\theta) = \mathbb{E}_{(x_{\text{text}},\, x_{\text{img}},\, x_{\text{aud}}) \sim \mathcal{D}}\Big[\, \mathcal{L}_{\text{text}}(\theta) + \mathcal{L}_{\text{img}}(\theta) + \mathcal{L}_{\text{aud}}(\theta) \,\Big]

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

Ljoint\mathcal{L}_{\text{joint}}

Symbol L_joint

LjL_joint is computed from the expected values combined on the right.

Understand this part →

θ\theta

Symbol θ

θ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

E(xtext, ximg, xaud)∼D\mathbb{E}_{(x_{\text{text}},\, x_{\text{img}},\, x_{\text{aud}}) \sim \mathcal{D}}

Symbol E_(x_text, x_img, x_aud) sim D

E_(xtx_text, xix_img, xax_aud) sim D appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

Ltext\mathcal{L}_{\text{text}}

Symbol L_text

LtL_text appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

Limg\mathcal{L}_{\text{img}}

Symbol L_img

LiL_img appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

Laud\mathcal{L}_{\text{aud}}

Symbol L_aud

LaL_aud appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

The documented limitation of native pretraining in this literature is not architectural instability — that problem belongs to the next section — but capability gaps that persist despite the joint training. Kosmos-1’s authors introduce a nonverbal Raven-style IQ test specifically to probe reasoning that is not mediated by language, and report that the model reaches only 26 percent accuracy against a random baseline of 17 percent, describing “a large performance gap between the current model and the average level of adults” even as they credit the model with demonstrating the basic capability [ 5 ] . Joint training from scratch buys deeper cross-modal integration; it does not, on this…
Read the full surrounding passage
The documented limitation of native pretraining in this literature is not architectural instability — that problem belongs to the next section — but capability gaps that persist despite the joint training. Kosmos-1’s authors introduce a nonverbal Raven-style IQ test specifically to probe reasoning that is not mediated by language, and report that the model reaches only 26 percent accuracy against a random baseline of 17 percent, describing “a large performance gap between the current model and the average level of adults” even as they credit the model with demonstrating the basic capability [ 5 ] . Joint training from scratch buys deeper cross-modal integration; it does not, on this evidence, buy strong abstract reasoning over visual patterns for free. All of θ\theta here is shared and randomly initialised at t=0 , and no term is held fixed while another is optimised. That is the formal content of “native”: contrast it with the modular recipe, where the vision and language terms are each minimised separately, over separate data, to convergence, long before the two are ever combined.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared

See this formula across 1 published context →

Browse the mathematical compendium →