← All parts of this equation

Equation 6 · Part 5 · Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared

Symbol L_img

Ljoint(θ)=E(xtext, ximg, xaud)∼D[ Ltext(θ)+Limg(θ)+Laud(θ) ]\mathcal{L}_{\text{joint}}(\theta) = \mathbb{E}_{(x_{\text{text}},\, x_{\text{img}},\, x_{\text{aud}}) \sim \mathcal{D}}\Big[\, \mathcal{L}_{\text{text}}(\theta) + \mathcal{L}_{\text{img}}(\theta) + \mathcal{L}_{\text{aud}}(\theta) \,\Big]
Limg\mathcal{L}_{\text{img}}

What this part means

LiL_img appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Its job in the formula

LiL_img appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

The passage around this formula

The documented limitation of native pretraining in this literature is not architectural instability — that problem belongs to the next section — but capability gaps that persist despite the joint training. Kosmos-1’s authors introduce a nonverbal Raven-style IQ test specifically to probe reasoning that is not mediated by language, and report that the model reaches only 26 percent accuracy against a random baseline of 17 percent, describing “a large performance gap between the current model and the average level of adults” even as they credit the model with demonstrating the basic capability [ 5 ] . Joint training from scratch buys deeper cross-modal integration; it does not, on this…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.