Foundation Models
Serving a Frontier Model: The KV Cache, Batching, and What a Token Actually Costs
Decoding is not limited by arithmetic. It is limited by how fast memory can be read, and almost every serving technique in production exists to work around that single fact.