Foundation Models
Adapters, Native Pretraining, and Unified Tokens: The Main Multimodal Architectures, Compared
Four research lineages solved the same problem — joining vision, audio and text into one system — with four different training recipes. Each recipe holds one cost fixed and lets another float, and the documented trade is an engineering choice, not a hierarchy.