← Mathematical compendium

Published equation contexts

Mresident≤BmemM_{\text{resident}} \le B_{\text{mem}}

Why this formula appears here

Compression pays mostly CreuseC_{\text{reuse}} , a documented small fraction of a from-scratch training cost [ 3 ] , but only if a suitable large model is available to reuse in the first place. Training small on purpose pays CcurateC_{\text{curate}} and CtrainC_{\text{train}} in full, with no discount, in exchange for a model that inherits nothing it was not deliberately given. Architecture search pays CsearchC_{\text{search}} , which can be enormous when paid fresh per target [ 7 ] or amortized across many targets when paid once as a supernet [ 10 ] , on top of whatever strategy trains the architecture it discovers. Sparse mixture-of-experts is the odd one out in this accounting: its saving shows up only in \bar…

Read the full article-specific guide →

Read the representative guide

MresidentM_{\text{resident}}

Symbol M_resident

MrM_resident is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
BmemB_{\text{mem}}

Symbol B_mem

BmB_mem is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Mresident≤BmemM_{\text{resident}} \le B_{\text{mem}}

Equation 39 · Edge AI & Electronics

Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions.

Compression pays mostly CreuseC_{\text{reuse}} , a documented small fraction of a from-scratch training cost [ 3 ] , but only if a suitable large model is available to reuse in the first place. Training small on purpose pays CcurateC_{\text{curate}} and CtrainC_{\text{train}} in full, with no discount, in exchange for a model that inherits nothing it was not deliberately given. Architecture search pays CsearchC_{\text{search}} , which can be enormous when paid fresh per target [ 7 ] or amortized across many targets when paid once as a supernet [ 10 ] , on top of whatever strategy trains the architecture it discovers. Sparse mixture-of-experts is the odd one out in this accounting: its saving shows up only in \bar…

Equation guide → · Article →