← All parts of this equation

Equation 12 · Part 1 · The Hardest Unsolved Problems in Small and On-Device AI

Symbol L

L(θ)=Lnew(θ)+λ∑iFi(θi−θi∗)2,\mathcal{L}(\theta) = \mathcal{L}_{\mathrm{new}}(\theta) + \lambda \sum_{i} F_i\left(\theta_i - \theta_i^{*}\right)^2,
L\mathcal{L}

What this part means

L is part of the quantity the equation computes from the expression on the right.

Its job in the formula

L is part of the quantity the equation computes from the expression on the right.

The passage around this formula

One standard mitigation is regularization: penalize the optimizer for moving parameters that mattered to earlier tasks. The best-known form estimates a per-parameter importance weight — commonly the diagonal of the Fisher information, FiF_i — from the old task, and adds it to the new loss: L(θ)=Lnew(θ)+λ∑iFi(θi−θi∗)2\mathcal{L}(\theta) = \mathcal{L}_{\mathrm{new}}(\theta) + \lambda \sum_{i} F_i\left(\theta_i - \theta_i^{*}\right)^2. where θi∗\theta_i^{*} is the old optimum for parameter i and λ\lambda sets how strongly the old task is protected. This equation exposes the actual trade rather than resolving it: raising λ\lambda protects old knowledge at the direct expense of how much the new update is allowed to change the model, and there is no value of λ\lambda that removes the trade — only one that relocates it. It also…

Read this part in the article →

Learn the underlying idea

A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.

Open the illustrated functions: inputs become outputs guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.