← Mathematical compendium

Published equation contexts

Clifetime=Creuse⏟compress  +  Ccurate+Ctrain⏟train small  +  Csearch⏟NAS  +  Q⋅cˉtok,withMresident≤BmemC_{\text{lifetime}} = \underbrace{C_{\text{reuse}}}_{\text{compress}} \;+\; \underbrace{C_{\text{curate}} + C_{\text{train}}}_{\text{train small}} \;+\; \underbrace{C_{\text{search}}}_{\text{NAS}} \;+\; Q \cdot \bar c_{\text{tok}}, \qquad \text{with} \quad M_{\text{resident}} \le B_{\text{mem}}

Why this formula appears here

Set side by side, the four strategies do not compete on a single scale, and a comparison that reduces them to one leaderboard number is not describing what any of them actually trades off. Each holds a different quantity fixed as “already spent” and treats a different quantity as the one still to be paid: Clifetime=Creuse⏟compress  +  Ccurate+Ctrain⏟train small  +  Csearch⏟NAS  +  Q⋅cˉtok,withMresident≤BmemC_{\text{lifetime}} = \underbrace{C_{\text{reuse}}}_{\text{compress}} \;+\; \underbrace{C_{\text{curate}} + C_{\text{train}}}_{\text{train small}} \;+\; \underbrace{C_{\text{search}}}_{\text{NAS}} \;+\; Q \cdot \bar c_{\text{tok}}, \qquad \text{with} \quad M_{\text{resident}} \le B_{\text{mem}}. Compression pays mostly CreuseC_{\text{reuse}} , a documented small fraction of a from-scratch training cost [ 3 ] , but only if a suitable large model is available to reuse in the first place. Training small on purpose pays CcurateC_{\text{curate}} and CtrainC_{\text{train}} in full, with no discount, in exchange for a model that inherits nothing it was not deliberately given.…

Read the full article-specific guide →

Read the representative guide

ClifetimeC_{\text{lifetime}}

Symbol C_lifetime

ClC_lifetime is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →
CtrainC_{\text{train}}

Symbol C_train

CtC_train is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →
cˉtok\bar c_{\text{tok}}

Symbol bar c_tok

bar ctc_tok is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →
MresidentM_{\text{resident}}

Symbol M_resident

MrM_resident is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →
BmemB_{\text{mem}}

Symbol B_mem

BmB_mem is one of the signed contributions combined to compute the quantity on the left.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Clifetime=Creuse⏟compress  +  Ccurate+Ctrain⏟train small  +  Csearch⏟NAS  +  Q⋅cˉtok,withMresident≤BmemC_{\text{lifetime}} = \underbrace{C_{\text{reuse}}}_{\text{compress}} \;+\; \underbrace{C_{\text{curate}} + C_{\text{train}}}_{\text{train small}} \;+\; \underbrace{C_{\text{search}}}_{\text{NAS}} \;+\; Q \cdot \bar c_{\text{tok}}, \qquad \text{with} \quad M_{\text{resident}} \le B_{\text{mem}}

Equation 33 · Edge AI & Electronics

Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions.

Set side by side, the four strategies do not compete on a single scale, and a comparison that reduces them to one leaderboard number is not describing what any of them actually trades off. Each holds a different quantity fixed as “already spent” and treats a different quantity as the one still to be paid: Clifetime=Creuse⏟compress  +  Ccurate+Ctrain⏟train small  +  Csearch⏟NAS  +  Q⋅cˉtok,withMresident≤BmemC_{\text{lifetime}} = \underbrace{C_{\text{reuse}}}_{\text{compress}} \;+\; \underbrace{C_{\text{curate}} + C_{\text{train}}}_{\text{train small}} \;+\; \underbrace{C_{\text{search}}}_{\text{NAS}} \;+\; Q \cdot \bar c_{\text{tok}}, \qquad \text{with} \quad M_{\text{resident}} \le B_{\text{mem}}. Compression pays mostly CreuseC_{\text{reuse}} , a documented small fraction of a from-scratch training cost [ 3 ] , but only if a suitable large model is available to reuse in the first place. Training small on purpose pays CcurateC_{\text{curate}} and CtrainC_{\text{train}} in full, with no discount, in exchange for a model that inherits nothing it was not deliberately given.…

Meanings in this article

Equation guide → · Article →