Equation 19 · Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol a^*
is part of the quantity the equation computes from the expression on the right.
Symbol a
a is part of the quantity the equation computes from the expression on the right.
Symbol A
a search space of candidate architectures designed in advance by the researchers, g(a) some measured deployment cost of architecture a.
Symbol L_val
al is one factor in the product that computes the quantity on the left.
Symbol w^*
is one factor in the product that computes the quantity on the left.
Symbol g
g is one factor in the product that computes the quantity on the left.
Symbol w
w is one factor in the product that computes the quantity on the left.
Symbol L_train
rain is one factor in the product that computes the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Zoph and Le established the modern form of the idea: a controller network, trained by reinforcement learning, proposes candidate child-network architectures, each of which is trained and evaluated, with the resulting performance used as a reward signal to improve the controller [ 7 ] . Formally, a NAS run of this kind is a constrained, nested optimization: . where is a search space of candidate architectures designed in advance by the researchers, g(a) some measured deployment cost of architecture a , and B a budget the target device imposes. Every term in that equation is a documented design choice, and the choice of g turns out to be where the edge-specific…
Read the full surrounding passage
Zoph and Le established the modern form of the idea: a controller network, trained by reinforcement learning, proposes candidate child-network architectures, each of which is trained and evaluated, with the resulting performance used as a reward signal to improve the controller [ 7 ] . Formally, a NAS run of this kind is a constrained, nested optimization: . where is a search space of candidate architectures designed in advance by the researchers, g(a) some measured deployment cost of architecture a , and B a budget the target device imposes. Every term in that equation is a documented design choice, and the choice of g turns out to be where the edge-specific literature does its most important work.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.