Equation 3 · One Model, Many Modalities: What Multimodal Systems Actually Share
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol W
W is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
for an image of height H and width W at patch size P . The quadratic relationship is the entire practical story of image tokenisation. Halving the patch size quadruples the sequence, and attention cost grows faster still. Every deployed system therefore resolves a three-way tension between input resolution, patch size, and context budget, and it resolves it by throwing away detail.
Sources cited in the article section
- [2] An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale ↗
- [3] Neural Discrete Representation Learning ↗
- [4] High Fidelity Neural Audio Compression ↗
- [5] Robust Speech Recognition via Large-Scale Weak Supervision ↗
These citations give research context. Read each source to check which claims it supports.
Return to One Model, Many Modalities: What Multimodal Systems Actually Share