Equation 33 · Part 6 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared
Symbol Q
What this part means
Q is one of the signed contributions combined to compute the quantity on the left.
Its job in the formula
Q is one of the signed contributions combined to compute the quantity on the left.
Full expression→Symbol Q→Article meaning
The passage around this formula
Set side by side, the four strategies do not compete on a single scale, and a comparison that reduces them to one leaderboard number is not describing what any of them actually trades off. Each holds a different quantity fixed as “already spent” and treats a different quantity as the one still to be paid: . Compression pays mostly , a documented small fraction of a from-scratch training cost [ 3 ] , but only if a suitable large model is available to reuse in the first place. Training small on purpose pays and in full, with no discount, in exchange for a model that inherits nothing it was not deliberately given.…
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the surrounding passage
- [3] Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning ↗
- [7] Neural Architecture Search with Reinforcement Learning ↗
- [10] Once for All: Train One Network and Specialize it for Efficient Deployment ↗
These citations provide research context; check each source for the exact claim it supports.