← Mathematical compendium

Published equation contexts

Mkv  ≈  2 L H dh b TM_{\mathrm{kv}} \;\approx\; 2 \, L \, H \, d_h \, b \, T

Why this formula appears here

None of the four sections above describes a system anyone ships in isolation. A contemporary AI accelerator composes several of these answers on top of each other, for a reason the earlier article in this series already established: capacity and bandwidth are separate constraints with separate symptoms, and a real workload — serving a large language model’s key-value cache is the canonical case — hits both at once. The KV cache for a single sequence grows as Mkv  ≈  2 L H dh b TM_{\mathrm{kv}} \;\approx\; 2 \, L \, H \, d_h \, b \, T. with L layers, H key/value heads, dhd_h head dimension, b bytes per stored element and T tokens of context — a quantity that scales linearly with context length and batch size, and has to be both stored somewhere and…

Read the full article-specific guide →

Read the representative guide

MkvM_{\mathrm{kv}}

Symbol M_kv

MkM_kv is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
HH

Symbol H

H is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
dhd_h

Symbol d_h

dhd_h is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
bb

Symbol b

b is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Its accuracy depends on the assumptions and range of use described in the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Mkv  ≈  2 L H dh b T,M_{\mathrm{kv}} \;\approx\; 2 \, L \, H \, d_h \, b \, T,

Equation 13 · Semiconductors

Comparing the Main Approaches to AI Memory Systems and the Bandwidth Wall

This equation gives an approximation: it relates the quantities while allowing an approximation.

None of the four sections above describes a system anyone ships in isolation. A contemporary AI accelerator composes several of these answers on top of each other, for a reason the earlier article in this series already established: capacity and bandwidth are separate constraints with separate symptoms, and a real workload — serving a large language model’s key-value cache is the canonical case — hits both at once. The KV cache for a single sequence grows as Mkv  ≈  2 L H dh b TM_{\mathrm{kv}} \;\approx\; 2 \, L \, H \, d_h \, b \, T. with L layers, H key/value heads, dhd_h head dimension, b bytes per stored element and T tokens of context — a quantity that scales linearly with context length and batch size, and has to be both stored somewhere and…

Meanings in this article

  • LL: the number of layers.
  • TT: the number of tokens.
Equation guide → · Article →