← All parts of this equation

Equation 6 · Part 7 · Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared

=

Ljoint(θ)=E(xtext, ximg, xaud)∼D[ Ltext(θ)+Limg(θ)+Laud(θ) ]\mathcal{L}_{\text{joint}}(\theta) = \mathbb{E}_{(x_{\text{text}},\, x_{\text{img}},\, x_{\text{aud}}) \sim \mathcal{D}}\Big[\, \mathcal{L}_{\text{text}}(\theta) + \mathcal{L}_{\text{img}}(\theta) + \mathcal{L}_{\text{aud}}(\theta) \,\Big]
=

What this part means

The expressions on both sides represent the same quantity under the stated assumptions.

Its job in the formula

The equals sign connects the complete expression on the left with the complete expression on the right. Both sides must have compatible units.

The passage around this formula

The documented limitation of native pretraining in this literature is not architectural instability — that problem belongs to the next section — but capability gaps that persist despite the joint training. Kosmos-1’s authors introduce a nonverbal Raven-style IQ test specifically to probe reasoning that is not mediated by language, and report that the model reaches only 26 percent accuracy against a random baseline of 17 percent, describing “a large performance gap between the current model and the average level of adults” even as they credit the model with demonstrating the basic capability [ 5 ] . Joint training from scratch buys deeper cross-modal integration; it does not, on this…

Read this part in the article →

Learn the underlying idea

An equals sign says that the expression on its left and the expression on its right have the same value under the stated definitions and assumptions.

Open the illustrated equality: what the equals sign claims guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.