Equation 12 · The Hardest Unsolved Problems in Small and On-Device AI
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol L
L is part of the quantity the equation computes from the expression on the right.
Symbol θ
θ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Symbol L_new
ew is one of the signed contributions combined to compute the quantity on the left.
Symbol λ
λ is one of the signed contributions combined to compute the quantity on the left.
Symbol i
i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Symbol theta_i
thet is one of the signed contributions combined to compute the quantity on the left.
Symbol theta_i^*
the old optimum for parameter i and sets how strongly the old task is protected.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subtraction
Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →Starting index or lower bound: i
This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
One standard mitigation is regularization: penalize the optimizer for moving parameters that mattered to earlier tasks. The best-known form estimates a per-parameter importance weight — commonly the diagonal of the Fisher information, — from the old task, and adds it to the new loss: . where is the old optimum for parameter i and sets how strongly the old task is protected. This equation exposes the actual trade rather than resolving it: raising protects old knowledge at the direct expense of how much the new update is allowed to change the model, and there is no value of that removes the trade — only one that relocates it. It also…
Read the full surrounding passage
One standard mitigation is regularization: penalize the optimizer for moving parameters that mattered to earlier tasks. The best-known form estimates a per-parameter importance weight — commonly the diagonal of the Fisher information, — from the old task, and adds it to the new loss: . where is the old optimum for parameter i and sets how strongly the old task is protected. This equation exposes the actual trade rather than resolving it: raising protects old knowledge at the direct expense of how much the new update is allowed to change the model, and there is no value of that removes the trade — only one that relocates it. It also exposes a cost specific to on-device deployment: computing and storing for every parameter, and doing so repeatedly as the device keeps learning, is itself memory and compute that a phone-class budget has to find room for, on top of whatever the update itself costs.
Sources cited in the article section
- [9] Continual Learning of Large Language Models: A Comprehensive Survey ↗
- [10] Online Continual Learning for Embedded Devices ↗
- [13] On-Device Language Models: A Comprehensive Review ↗
- [11] Catastrophic Forgetting in LLMs: A Comparative Analysis Across Language Tasks ↗
These citations give research context. Read each source to check which claims it supports.
Return to The Hardest Unsolved Problems in Small and On-Device AI