Equation 11 · The Hardest Unsolved Problems in Small and On-Device AI
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the computing and storing. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
One standard mitigation is regularization: penalize the optimizer for moving parameters that mattered to earlier tasks. The best-known form estimates a per-parameter importance weight — commonly the diagonal of the Fisher information, — from the old task, and adds it to the new loss:
Sources cited in the article section
- [9] Continual Learning of Large Language Models: A Comprehensive Survey ↗
- [10] Online Continual Learning for Embedded Devices ↗
- [13] On-Device Language Models: A Comprehensive Review ↗
- [11] Catastrophic Forgetting in LLMs: A Comparative Analysis Across Language Tasks ↗
These citations give research context. Read each source to check which claims it supports.
Return to The Hardest Unsolved Problems in Small and On-Device AI