Equation 7 · Part 2 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared
subscript
subscript
What this part means
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
Its job in the formula
A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.
Full expression→subscript→Article meaning
The passage around this formula
The strategy carries two structural limitations that follow directly from how it works, not from any one paper’s shortcoming. First, it requires access to the large model itself, or at minimum to its outputs at sufficient fidelity to train against — a dependency none of the other three strategies share. If the frontier model is API-only, gated, or simply unavailable to the team doing the compressing, the achievable fidelity of “dark knowledge” transfer is bounded by whatever access is actually granted. Second, and more fundamentally, a compressed model is trained to match its teacher, not trained against the underlying task from scratch — whatever the teacher gets systematically wrong,…
Learn the underlying idea
A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.
Open the illustrated subscripts: which member of a family? guide →
Sources cited in the article section
- [1] Distilling the Knowledge in a Neural Network ↗
- [3] Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning ↗
- [2] Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding ↗
These citations provide research context; check each source for the exact claim it supports.