Symbol M_KV
V is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
Fact. The memory a transformer must hold for one in-flight generation — its key-value cache — grows linearly in exactly the variables that make serving expensive: sequence length and batch size. For a model with L layers, key-value heads, head dimension , sequence length S , batch size B , and p bytes stored per element, the cache occupies approximately . bytes, the leading factor of two accounting for keys and values together. That equation is worth writing out because it exposes the one real lever every mitigation below actually pulls: none of them escapes linear growth in S and B . Each instead shrinks one of the other factors — most consequentially…
V is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →v is an input to the expression that computes the quantity on the left.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 13 · Semiconductors
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
Fact. The memory a transformer must hold for one in-flight generation — its key-value cache — grows linearly in exactly the variables that make serving expensive: sequence length and batch size. For a model with L layers, key-value heads, head dimension , sequence length S , batch size B , and p bytes stored per element, the cache occupies approximately . bytes, the leading factor of two accounting for keys and values together. That equation is worth writing out because it exposes the one real lever every mitigation below actually pulls: none of them escapes linear growth in S and B . Each instead shrinks one of the other factors — most consequentially…