← All parts of this equation

Equation 15 · Part 1 · How AI Inference Serving Actually Works

Symbol b

bb
bb

What this part means

the number of concurrently served sequences.

Its job in the formula

b is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Where the article explains it

For a model with L layers, hkvh_{kv} key-value attention heads, head dimension dhd_h , holding a sequence of length s across b concurrently served sequences, stored at p bytes per element, the total resident size is MkvM_{\mathrm{kv}} = 2\,L\,hkvh_{kv}\,dhd_h\,s\,b\,p with the factor of two accounting for storing both keys and values.

The passage around this formula

…from this equation, and both matter more than the equation’s arithmetic itself. It scales linearly with context length, so a conversation twice as long holds twice the cache. And it scales linearly with the batch b — the exact quantity the previous section identified as the only lever available to make a memory-bound decode step efficient. The key-value cache therefore competes directly, in the same pool of device memory, with the batch size continuous batching is trying to grow. It is not a side cost of…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.