← Back to article

Equation 12 · The Hardest Unsolved Problems in Small and On-Device AI

What does this equation mean?

L(θ)=Lnew(θ)+λ∑iFi(θi−θi∗)2,\mathcal{L}(\theta) = \mathcal{L}_{\mathrm{new}}(\theta) + \lambda \sum_{i} F_i\left(\theta_i - \theta_i^{*}\right)^2,

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationsL_new(θ) + λ sum_i F_i(theta_i - theta_i^*)^2
Result or conditionL(θ)
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

L\mathcal{L}

Symbol L

L is part of the quantity the equation computes from the expression on the right.

Understand this part →

θ\theta

Symbol θ

θ is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

Lnew\mathcal{L}_{\mathrm{new}}

Symbol L_new

LnL_new is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

λ\lambda

Symbol λ

λ is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

ii

Symbol i

i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Understand this part →

FiF_i

Symbol F_i

the computing and storing.

Understand this part →

θi\theta_i

Symbol theta_i

thetaia_i is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

θi∗\theta_i^{*}

Symbol theta_i^*

the old optimum for parameter i and λ\lambda sets how strongly the old task is protected.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subtraction

subtraction

Subtract the following term or group from the preceding one. A leading minus marks a negative quantity.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

superscript

superscript

A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.

Understand this part →

See an illustrated explanation →
ii

Starting index or lower bound: i

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

One standard mitigation is regularization: penalize the optimizer for moving parameters that mattered to earlier tasks. The best-known form estimates a per-parameter importance weight — commonly the diagonal of the Fisher information, FiF_i — from the old task, and adds it to the new loss: L(θ)=Lnew(θ)+λ∑iFi(θi−θi∗)2\mathcal{L}(\theta) = \mathcal{L}_{\mathrm{new}}(\theta) + \lambda \sum_{i} F_i\left(\theta_i - \theta_i^{*}\right)^2. where θi∗\theta_i^{*} is the old optimum for parameter i and λ\lambda sets how strongly the old task is protected. This equation exposes the actual trade rather than resolving it: raising λ\lambda protects old knowledge at the direct expense of how much the new update is allowed to change the model, and there is no value of λ\lambda that removes the trade — only one that relocates it. It also…
Read the full surrounding passage
One standard mitigation is regularization: penalize the optimizer for moving parameters that mattered to earlier tasks. The best-known form estimates a per-parameter importance weight — commonly the diagonal of the Fisher information, FiF_i — from the old task, and adds it to the new loss: L(θ)=Lnew(θ)+λ∑iFi(θi−θi∗)2\mathcal{L}(\theta) = \mathcal{L}_{\mathrm{new}}(\theta) + \lambda \sum_{i} F_i\left(\theta_i - \theta_i^{*}\right)^2. where θi∗\theta_i^{*} is the old optimum for parameter i and λ\lambda sets how strongly the old task is protected. This equation exposes the actual trade rather than resolving it: raising λ\lambda protects old knowledge at the direct expense of how much the new update is allowed to change the model, and there is no value of λ\lambda that removes the trade — only one that relocates it. It also exposes a cost specific to on-device deployment: computing and storing FiF_i for every parameter, and doing so repeatedly as the device keeps learning, is itself memory and compute that a phone-class budget has to find room for, on top of whatever the update itself costs.

Read the equation in its article →

Sources cited in the article section

These citations give research context. Read each source to check which claims it supports.

Return to The Hardest Unsolved Problems in Small and On-Device AI

See this formula across 1 published context →

Browse the mathematical compendium →