Symbol M_KV
V is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
The key-value cache is the part of a serving system’s memory footprint that a purchasing decision cannot fix, because it does not exist until a request is running. It grows with every generated token, in proportion to context length, batch size, layer count and head dimension, and its lifetime is exactly the lifetime of the request that owns it — arriving and departing unpredictably, at whatever size that request’s conversation happens to reach. A capacity plan sized for the average context length will be wrong for the tail, and a naive contiguous allocator sized for the tail will waste most of its memory on requests that never reach it. with L transformer layers, H key-value heads, head…
V is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →L is one factor in the product that computes the quantity on the left.
Read this term in its guide →H is one factor in the product that computes the quantity on the left.
Read this term in its guide →is one factor in the product that computes the quantity on the left.
Read this term in its guide →b is one factor in the product that computes the quantity on the left.
Read this term in its guide →s is one factor in the product that computes the quantity on the left.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 1 · Semiconductors
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
The key-value cache is the part of a serving system’s memory footprint that a purchasing decision cannot fix, because it does not exist until a request is running. It grows with every generated token, in proportion to context length, batch size, layer count and head dimension, and its lifetime is exactly the lifetime of the request that owns it — arriving and departing unpredictably, at whatever size that request’s conversation happens to reach. A capacity plan sized for the average context length will be wrong for the tail, and a naive contiguous allocator sized for the tail will waste most of its memory on requests that never reach it. with L transformer layers, H key-value heads, head…