← Back to article

Equation 18 · How AI Datacenter Systems Engineering Actually Works

What does this equation mean?

CC

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

the wall-clock cost of writing one checkpoint and M be the job’s mean time between failures. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

CC

Symbol C

the wall-clock cost of writing one checkpoint and M be the job’s mean time between failures.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Meta’s account of training the 405-billion-parameter Llama 3 model gives the concrete numbers behind all of this. Each GPU’s checkpointed state runs one to four gigabytes, and the stated design goals are exactly C and T : “minimizing GPU pause time during checkpointing and increasing checkpoint frequency to reduce the amount of lost work after a recovery” [ 9 ] . Over a 54-day monitoring window the run absorbed 466 total job interruptions — 47 planned and 419 unexpected, of which roughly 78% were confirmed or suspected hardware issues and 58.7% were GPU-related specifically — while still achieving “higher than 90% effective training time,” with only three interruptions requiring significant…
Read the full surrounding passage
Meta’s account of training the 405-billion-parameter Llama 3 model gives the concrete numbers behind all of this. Each GPU’s checkpointed state runs one to four gigabytes, and the stated design goals are exactly C and T : “minimizing GPU pause time during checkpointing and increasing checkpoint frequency to reduce the amount of lost work after a recovery” [ 9 ] . Over a 54-day monitoring window the run absorbed 466 total job interruptions — 47 planned and 419 unexpected, of which roughly 78% were confirmed or suspected hardware issues and 58.7% were GPU-related specifically — while still achieving “higher than 90% effective training time,” with only three interruptions requiring significant manual intervention and the rest handled by automation [ 9 ] . That is the checkpoint-frequency tradeoff resolved in production: at that failure rate, the interval had to be short enough that losing an average interval’s worth of work, repeatedly, still left ninety cents of every dollar of GPU-time doing useful computation.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to How AI Datacenter Systems Engineering Actually Works

Browse the mathematical compendium →