← All parts of this equation

Equation 18 · Part 1 · How AI Inference Serving Actually Works

Symbol E

E[tokens per round]=1−αγ+11−α.\mathbb{E}[\text{tokens per round}] = \frac{1-\alpha^{\gamma+1}}{1-\alpha}.
E\mathbb{E}

What this part means

The expected value operator: the probability-weighted average of the quantity inside its brackets.

Its job in the formula

E is part of the quantity the equation computes from the expression on the right.

The passage around this formula

The size of the win has a clean shape. Model the draft’s acceptance probability as α\alpha per token, roughly constant and independent across the γ\gamma tokens proposed in a round — an idealization real traffic does not fully satisfy, but a useful one for seeing the ceiling. The expected number of tokens accepted per verification round is then E[tokens per round]=1−αγ+11−α\mathbb{E}[\text{tokens per round}] = \frac{1-\alpha^{\gamma+1}}{1-\alpha}. This rises with both α\alpha and γ\gamma , but with steeply diminishing returns in γ\gamma for any α\alpha below one: drafting fifty tokens ahead does not buy anywhere near fifty accepted tokens, because the marginal proposal deep into a long draft is unlikely to be exactly what the target would have generated. The bottleneck decode…

Read this part in the article →

Learn the underlying idea

Probability assigns a number from 0 to 1 to an event under a stated model. Zero means impossible within that model; one means certain.

Open the illustrated probability: a quantified chance guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.