← Mathematical compendium

Published equation contexts

MKV=2⋅L⋅H⋅dh⋅b⋅s⋅pM_{\mathrm{KV}} = 2 \cdot L \cdot H \cdot d_h \cdot b \cdot s \cdot p

Why this formula appears here

The key-value cache is the part of a serving system’s memory footprint that a purchasing decision cannot fix, because it does not exist until a request is running. It grows with every generated token, in proportion to context length, batch size, layer count and head dimension, and its lifetime is exactly the lifetime of the request that owns it — arriving and departing unpredictably, at whatever size that request’s conversation happens to reach. A capacity plan sized for the average context length will be wrong for the tail, and a naive contiguous allocator sized for the tail will waste most of its memory on requests that never reach it. with L transformer layers, H key-value heads, dhd_h head…

Read the full article-specific guide →

Read the representative guide

MKVM_{\mathrm{KV}}

Symbol M_KV

MKM_KV is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

MKV=2⋅L⋅H⋅dh⋅b⋅s⋅pM_{\mathrm{KV}} = 2 \cdot L \cdot H \cdot d_h \cdot b \cdot s \cdot p

Equation 1 · Semiconductors

AI Memory Systems and the Bandwidth Wall in Practice: An Advanced Technical Guide

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

The key-value cache is the part of a serving system’s memory footprint that a purchasing decision cannot fix, because it does not exist until a request is running. It grows with every generated token, in proportion to context length, batch size, layer count and head dimension, and its lifetime is exactly the lifetime of the request that owns it — arriving and departing unpredictably, at whatever size that request’s conversation happens to reach. A capacity plan sized for the average context length will be wrong for the tail, and a naive contiguous allocator sized for the tail will waste most of its memory on requests that never reach it. with L transformer layers, H key-value heads, dhd_h head…

Meanings in this article

  • pp: the number of bytes.
Equation guide → · Article →