← Mathematical compendium

Published equation contexts

MKV  =  2 n h d e b lM_{\mathrm{KV}} \;=\; 2 \, n \, h \, d \, e \, b \, l

Why this formula appears here

Hooper and colleagues, working on KV cache compression for very long contexts, state the resulting footprint precisely: for a model with n layers and h attention heads of dimension d , stored using e bytes per element, the KV cache size for batch size b and sequence length l is MKV  =  2 n h d e b lM_{\mathrm{KV}} \;=\; 2 \, n \, h \, d \, e \, b \, l . which grows linearly in both batch size and sequence length, with the leading factor of two accounting for storing both keys and values [ 9 ] . That single equation is the whole mechanism: nothing about it is a design choice an inference engineer can simply decline. Extend the conversation, and l grows; serve more requests at once, and b grows; either way MKVM_{\mathrm{KV}} grows with it, and Hooper…

Read the full article-specific guide →

Read the representative guide

MKVM_{\mathrm{KV}}

Symbol M_KV

MKM_KV is part of the quantity the equation computes from the expression on the right.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

MKV  =  2 n h d e b l,M_{\mathrm{KV}} \;=\; 2 \, n \, h \, d \, e \, b \, l ,

Equation 12 · Semiconductors

How AI Memory Systems and the Bandwidth Wall Actually Work

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Hooper and colleagues, working on KV cache compression for very long contexts, state the resulting footprint precisely: for a model with n layers and h attention heads of dimension d , stored using e bytes per element, the KV cache size for batch size b and sequence length l is MKV  =  2 n h d e b lM_{\mathrm{KV}} \;=\; 2 \, n \, h \, d \, e \, b \, l . which grows linearly in both batch size and sequence length, with the leading factor of two accounting for storing both keys and values [ 9 ] . That single equation is the whole mechanism: nothing about it is a design choice an inference engineer can simply decline. Extend the conversation, and l grows; serve more requests at once, and b grows; either way MKVM_{\mathrm{KV}} grows with it, and Hooper…

Meanings in this article

  • nn: the number of layers.
  • ee: the bytes stored per element.
Equation guide → · Article →