Equation 4 · Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol theta_vision
thetision is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
This is the formal shape of “freeze the backbones”: the gradient update is masked to zero everywhere except the bridge parameters , so || — a few hundred million parameters in BLIP-2’s case, dramatically fewer than either tower — is the entire quantity being optimised, and and contribute a forward pass but never a gradient.
Sources cited in the article section
- [3] Flamingo: a Visual Language Model for Few-Shot Learning ↗
- [2] BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models ↗
- [1] Visual Instruction Tuning ↗
These citations give research context. Read each source to check which claims it supports.