Symbol M_kv
v is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Published equation contexts
None of the four sections above describes a system anyone ships in isolation. A contemporary AI accelerator composes several of these answers on top of each other, for a reason the earlier article in this series already established: capacity and bandwidth are separate constraints with separate symptoms, and a real workload — serving a large language model’s key-value cache is the canonical case — hits both at once. The KV cache for a single sequence grows as . with L layers, H key/value heads, head dimension, b bytes per stored element and T tokens of context — a quantity that scales linearly with context length and batch size, and has to be both stored somewhere and…
v is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →H is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →b is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Its accuracy depends on the assumptions and range of use described in the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 13 · Semiconductors
This equation gives an approximation: it relates the quantities while allowing an approximation.
None of the four sections above describes a system anyone ships in isolation. A contemporary AI accelerator composes several of these answers on top of each other, for a reason the earlier article in this series already established: capacity and bandwidth are separate constraints with separate symptoms, and a real workload — serving a large language model’s key-value cache is the canonical case — hits both at once. The KV cache for a single sequence grows as . with L layers, H key/value heads, head dimension, b bytes per stored element and T tokens of context — a quantity that scales linearly with context length and batch size, and has to be both stored somewhere and…