Symbol M_kv
v is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Published equation contexts
A transformer decoder generates autoregressively: each new token attends over all previous positions [ 1 ] . To avoid recomputing the whole prefix at every step, implementations retain the per-layer key and value tensors for every past position. That is the key–value cache, and its size grows linearly with sequence length: . for L layers, key–value heads, head dimension , sequence length s , batch size b , and p bytes per element. The factor of two counts keys and values.
v is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →v is an input to the expression that computes the quantity on the left.
Read this term in its guide →p is an input to the expression that computes the quantity on the left.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 1 · Foundation Models
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
A transformer decoder generates autoregressively: each new token attends over all previous positions [ 1 ] . To avoid recomputing the whole prefix at every step, implementations retain the per-layer key and value tensors for every past position. That is the key–value cache, and its size grows linearly with sequence length: . for L layers, key–value heads, head dimension , sequence length s , batch size b , and p bytes per element. The factor of two counts keys and values.