← Mathematical compendium

Published equation contexts

Ljoint(θ)=E(xtext, ximg, xaud)∼D[ Ltext(θ)+Limg(θ)+Laud(θ) ]\mathcal{L}_{\text{joint}}(\theta) = \mathbb{E}_{(x_{\text{text}},\, x_{\text{img}},\, x_{\text{aud}}) \sim \mathcal{D}}\Big[\, \mathcal{L}_{\text{text}}(\theta) + \mathcal{L}_{\text{img}}(\theta) + \mathcal{L}_{\text{aud}}(\theta) \,\Big]

Why this formula appears here

The documented limitation of native pretraining in this literature is not architectural instability — that problem belongs to the next section — but capability gaps that persist despite the joint training. Kosmos-1’s authors introduce a nonverbal Raven-style IQ test specifically to probe reasoning that is not mediated by language, and report that the model reaches only 26 percent accuracy against a random baseline of 17 percent, describing “a large performance gap between the current model and the average level of adults” even as they credit the model with demonstrating the basic capability [ 5 ] . Joint training from scratch buys deeper cross-modal integration; it does not, on this…

Read the full article-specific guide →

Read the representative guide

Ljoint\mathcal{L}_{\text{joint}}

Symbol L_joint

LjL_joint is computed from the expected values combined on the right.

Read this term in its guide →
θ\theta

Symbol θ

θ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Read this term in its guide →
E(xtext, ximg, xaud)∼D\mathbb{E}_{(x_{\text{text}},\, x_{\text{img}},\, x_{\text{aud}}) \sim \mathcal{D}}

Symbol E_(x_text, x_img, x_aud) sim D

E_(xtx_text, xix_img, xax_aud) sim D appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
Ltext\mathcal{L}_{\text{text}}

Symbol L_text

LtL_text appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
Limg\mathcal{L}_{\text{img}}

Symbol L_img

LiL_img appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →
Laud\mathcal{L}_{\text{aud}}

Symbol L_aud

LaL_aud appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Ljoint(θ)=E(xtext, ximg, xaud)∼D[ Ltext(θ)+Limg(θ)+Laud(θ) ]\mathcal{L}_{\text{joint}}(\theta) = \mathbb{E}_{(x_{\text{text}},\, x_{\text{img}},\, x_{\text{aud}}) \sim \mathcal{D}}\Big[\, \mathcal{L}_{\text{text}}(\theta) + \mathcal{L}_{\text{img}}(\theta) + \mathcal{L}_{\text{aud}}(\theta) \,\Big]

Equation 6 · Foundation Models

Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

The documented limitation of native pretraining in this literature is not architectural instability — that problem belongs to the next section — but capability gaps that persist despite the joint training. Kosmos-1’s authors introduce a nonverbal Raven-style IQ test specifically to probe reasoning that is not mediated by language, and report that the model reaches only 26 percent accuracy against a random baseline of 17 percent, describing “a large performance gap between the current model and the average level of adults” even as they credit the model with demonstrating the basic capability [ 5 ] . Joint training from scratch buys deeper cross-modal integration; it does not, on this…

Equation guide → · Article →