Equation 6 · Part 3 · Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared
Symbol E_(x_text, x_img, x_aud) sim D
What this part means
E_(ext, mg, ud) sim D appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Its job in the formula
E_(ext, mg, ud) sim D appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Full expression→Symbol E_(x_text, x_img, x_aud) sim D→Article meaning
The passage around this formula
The documented limitation of native pretraining in this literature is not architectural instability — that problem belongs to the next section — but capability gaps that persist despite the joint training. Kosmos-1’s authors introduce a nonverbal Raven-style IQ test specifically to probe reasoning that is not mediated by language, and report that the model reaches only 26 percent accuracy against a random baseline of 17 percent, describing “a large performance gap between the current model and the average level of adults” even as they credit the model with demonstrating the basic capability [ 5 ] . Joint training from scratch buys deeper cross-modal integration; it does not, on this…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
See this notation across published equations →
Sources cited in the surrounding passage
These citations provide research context; check each source for the exact claim it supports.