← All parts of this equation

Equation 6 · Part 2 · Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared

Symbol θ

Ljoint(θ)=E(xtext, ximg, xaud)∼D[ Ltext(θ)+Limg(θ)+Laud(θ) ]\mathcal{L}_{\text{joint}}(\theta) = \mathbb{E}_{(x_{\text{text}},\, x_{\text{img}},\, x_{\text{aud}}) \sim \mathcal{D}}\Big[\, \mathcal{L}_{\text{text}}(\theta) + \mathcal{L}_{\text{img}}(\theta) + \mathcal{L}_{\text{aud}}(\theta) \,\Big]
θ\theta

What this part means

θ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Its job in the formula

θ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

The passage around this formula

…model with demonstrating the basic capability [ 5 ] . Joint training from scratch buys deeper cross-modal integration; it does not, on this evidence, buy strong abstract reasoning over visual patterns for free. All of θ\theta here is shared and randomly initialised at t=0 , and no term is held fixed while another is optimised. That is the formal content of “native”: contrast it with the modular recipe, where the vision and language terms are each minimised separately, over separate data, to convergence, long…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.