Equation 6 · Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol L_joint
oint is computed from the expected values combined on the right.
Symbol θ
θ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Symbol E_(x_text, x_img, x_aud) sim D
E_(ext, mg, ud) sim D appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol L_text
ext appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol L_img
mg appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
Symbol L_aud
ud appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The documented limitation of native pretraining in this literature is not architectural instability — that problem belongs to the next section — but capability gaps that persist despite the joint training. Kosmos-1’s authors introduce a nonverbal Raven-style IQ test specifically to probe reasoning that is not mediated by language, and report that the model reaches only 26 percent accuracy against a random baseline of 17 percent, describing “a large performance gap between the current model and the average level of adults” even as they credit the model with demonstrating the basic capability [ 5 ] . Joint training from scratch buys deeper cross-modal integration; it does not, on this…
Read the full surrounding passage
The documented limitation of native pretraining in this literature is not architectural instability — that problem belongs to the next section — but capability gaps that persist despite the joint training. Kosmos-1’s authors introduce a nonverbal Raven-style IQ test specifically to probe reasoning that is not mediated by language, and report that the model reaches only 26 percent accuracy against a random baseline of 17 percent, describing “a large performance gap between the current model and the average level of adults” even as they credit the model with demonstrating the basic capability [ 5 ] . Joint training from scratch buys deeper cross-modal integration; it does not, on this evidence, buy strong abstract reasoning over visual patterns for free. All of here is shared and randomly initialised at t=0 , and no term is held fixed while another is optimised. That is the formal content of “native”: contrast it with the modular recipe, where the vision and language terms are each minimised separately, over separate data, to convergence, long before the two are ever combined.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.