← All parts of this equation

Equation 6 · Part 1 · Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared

Symbol L_joint

Ljoint(θ)=E(xtext, ximg, xaud)∼D[ Ltext(θ)+Limg(θ)+Laud(θ) ]\mathcal{L}_{\text{joint}}(\theta) = \mathbb{E}_{(x_{\text{text}},\, x_{\text{img}},\, x_{\text{aud}}) \sim \mathcal{D}}\Big[\, \mathcal{L}_{\text{text}}(\theta) + \mathcal{L}_{\text{img}}(\theta) + \mathcal{L}_{\text{aud}}(\theta) \,\Big]
Ljoint\mathcal{L}_{\text{joint}}

What this part means

LjL_joint is computed from the expected values combined on the right.

Its job in the formula

LjL_joint is computed from the expected values combined on the right.

The passage around this formula

The documented limitation of native pretraining in this literature is not architectural instability — that problem belongs to the next section — but capability gaps that persist despite the joint training. Kosmos-1’s authors introduce a nonverbal Raven-style IQ test specifically to probe reasoning that is not mediated by language, and report that the model reaches only 26 percent accuracy against a random baseline of 17 percent, describing “a large performance gap between the current model and the average level of adults” even as they credit the model with demonstrating the basic capability [ 5 ] . Joint training from scratch buys deeper cross-modal integration; it does not, on this…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.