← Back to article

Equation 25 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

What does this equation mean?

g(a)g(a)

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

gg

Symbol g

g is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

aa

Symbol a

a is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

The first response changes what g(a) measures. Tan and colleagues built MnasNet’s rationale explicitly around the observation that operation counts are a poor stand-in for real efficiency, so the search “directly measures real-world inference latency by executing the model on mobile phones,” and reported an architecture 1.8x faster than MobileNetV2 at 0.5% higher accuracy, and 2.3x faster than NASNet at 1.2% higher accuracy [ 8 ] . That is a harder objective to satisfy than a proxy count, because it cannot be gamed by an architecture that is cheap on paper and slow in practice. The second response changes how often the expensive part of the search has to be paid at all. Cai and colleagues…
Read the full surrounding passage
The first response changes what g(a) measures. Tan and colleagues built MnasNet’s rationale explicitly around the observation that operation counts are a poor stand-in for real efficiency, so the search “directly measures real-world inference latency by executing the model on mobile phones,” and reported an architecture 1.8x faster than MobileNetV2 at 0.5% higher accuracy, and 2.3x faster than NASNet at 1.2% higher accuracy [ 8 ] . That is a harder objective to satisfy than a proxy count, because it cannot be gamed by an architecture that is cheap on paper and slow in practice. The second response changes how often the expensive part of the search has to be paid at all. Cai and colleagues named the underlying problem directly — that manually designing or running NAS “for each case” of a target device “is computationally prohibitive” — and instead trained one large “supernet” a single time, from which a specialized sub-network for a given device could be extracted “without additional training” [ 10 ] . That converts a cost that would otherwise recur for every device generation into a cost paid once, followed by cheap selection.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared

See this formula across 2 published contexts →

Browse the mathematical compendium →