← All parts of this equation

Equation 18 · Part 1 · Comparing the Main Approaches to AI Memory Systems and the Bandwidth Wall

Symbol T

TT
TT

What this part means

the number of tokens.

Its job in the formula

T is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Where the article explains it

with L layers, H key/value heads, dhd_h head dimension, b bytes per stored element and T tokens of context — a quantity that scales linearly with context length and batch size, and has to be both stored somewhere and streamed through the arithmetic units on every generated token.

The passage around this formula

with L layers, H key/value heads, dhd_h head dimension, b bytes per stored element and T tokens of context — a quantity that scales linearly with context length and batch size, and has to be both stored somewhere and streamed through the arithmetic units on every generated token. Pope and colleagues showed that partitioning strategies which shrink the key/value representation, such…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.