← Mathematical compendium

Published equation contexts

τopt≈2 δ M\tau_{\mathrm{opt}} \approx \sqrt{2\,\delta\,M}

Why this formula appears here

If a training job cannot simply be re-run task-by-task the way a MapReduce job could, recovering from a failure depends entirely on how recently and how cheaply its state was saved. Checkpointing research therefore had to become a first-class systems problem in its own right rather than a background convenience, and its mathematics is old. A 2024 re-derivation of the classical result on checkpoint scheduling shows that the loss-minimizing interval between checkpoints is proportional to the square root of the product of checkpoint save time and mean time to failure, τopt≈2 δ M\tau_{\mathrm{opt}} \approx \sqrt{2\,\delta\,M}. with δ\delta the time cost of writing one checkpoint and M the mean time between failures — and the paper…

Read the full article-specific guide →

Read the representative guide

τopt\tau_{\mathrm{opt}}

Symbol tau_opt

tauou_opt is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
δ\delta

Symbol delta

delta is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
MM

Symbol M

M is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Its accuracy depends on the assumptions and range of use described in the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

τopt≈2 δ M,\tau_{\mathrm{opt}} \approx \sqrt{2\,\delta\,M},

Equation 1 · Datacenters

From Origins to Frontier: A History of AI Datacenter Systems Engineering

This equation gives an approximation: it relates the quantities while allowing an approximation.

If a training job cannot simply be re-run task-by-task the way a MapReduce job could, recovering from a failure depends entirely on how recently and how cheaply its state was saved. Checkpointing research therefore had to become a first-class systems problem in its own right rather than a background convenience, and its mathematics is old. A 2024 re-derivation of the classical result on checkpoint scheduling shows that the loss-minimizing interval between checkpoints is proportional to the square root of the product of checkpoint save time and mean time to failure, τopt≈2 δ M\tau_{\mathrm{opt}} \approx \sqrt{2\,\delta\,M}. with δ\delta the time cost of writing one checkpoint and M the mean time between failures — and the paper…

Equation guide → · Article →