Equation 13 · The Token Tax of Giving a Model More Tools
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the shortlist shown to the model. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
subscript
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
A separate, applied study on enterprise agent routing gives a way to decompose exactly what breaks as a real tool catalogue scales, rather than only that something does. Scaling from 10 to 110 candidate agents or tools, the authors report routing F1 on under-specified requests dropping 16 to 23 percentage points across the models tested, and split that drop with an oracle analysis into two distinct components: a retrieval gap , the model’s failure to surface the correct tool at all, and a confusion gap , a roughly 10-percentage-point reduction in the theoretical best-case score that persists even when retrieval is assumed perfect [ 11 ] . Written as a decomposition of the error a shortlist…
Read the full surrounding passage
A separate, applied study on enterprise agent routing gives a way to decompose exactly what breaks as a real tool catalogue scales, rather than only that something does. Scaling from 10 to 110 candidate agents or tools, the authors report routing F1 on under-specified requests dropping 16 to 23 percentage points across the models tested, and split that drop with an oracle analysis into two distinct components: a retrieval gap , the model’s failure to surface the correct tool at all, and a confusion gap , a roughly 10-percentage-point reduction in the theoretical best-case score that persists even when retrieval is assumed perfect [ 11 ] . Written as a decomposition of the error a shortlist of size m drawn from a full catalogue of n tools produces, where is the one the request actually calls for and is the shortlist shown to the model:
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.