Equation 37 · Part 2 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared
subscript
subscript
What this part means
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
Its job in the formula
A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.
Full expression→subscript→Article meaning
The passage around this formula
Compression pays mostly , a documented small fraction of a from-scratch training cost [ 3 ] , but only if a suitable large model is available to reuse in the first place. Training small on purpose pays and in full, with no discount, in exchange for a model that inherits nothing it was not deliberately given. Architecture search pays , which can be enormous when paid fresh per target [ 7 ] or amortized across many targets when paid once as a supernet [ 10 ] , on top of whatever strategy trains the architecture it discovers. Sparse mixture-of-experts is the odd one out in this accounting: its saving shows up only in \bar…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
Sources cited in the surrounding passage
- [3] Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning ↗
- [7] Neural Architecture Search with Reinforcement Learning ↗
- [10] Once for All: Train One Network and Specialize it for Efficient Deployment ↗
These citations provide research context; check each source for the exact claim it supports.