← Back to article

Equation 14 · The Token Tax of Giving a Model More Tools

What does this equation mean?

error(m)  =  Pr⁡[tool⋆∉Sm]⏟retrieval gap, shrinks as m grows  +  Pr⁡[miscall∣tool⋆∈Sm]⏟confusion gap, grows as m grows\mathrm{error}(m) \;=\; \underbrace{\Pr[\mathrm{tool}^\star \notin S_m]}_{\text{retrieval gap, shrinks as } m \text{ grows}} \;+\; \underbrace{\Pr[\text{miscall} \mid \mathrm{tool}^\star \in S_m]}_{\text{confusion gap, grows as } m \text{ grows}}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

mm

Symbol m

m is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

SmS_m

Symbol S_m

the shortlist shown to the model: [displayed formula].

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

Pr⁡\Pr

Probability operator

The probability operator gives the chance of the event named inside its brackets or parentheses.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

A separate, applied study on enterprise agent routing gives a way to decompose exactly what breaks as a real tool catalogue scales, rather than only that something does. Scaling from 10 to 110 candidate agents or tools, the authors report routing F1 on under-specified requests dropping 16 to 23 percentage points across the models tested, and split that drop with an oracle analysis into two distinct components: a retrieval gap , the model’s failure to surface the correct tool at all, and a confusion gap , a roughly 10-percentage-point reduction in the theoretical best-case score that persists even when retrieval is assumed perfect [ 11 ] . Written as a decomposition of the error a shortlist…
Read the full surrounding passage
A separate, applied study on enterprise agent routing gives a way to decompose exactly what breaks as a real tool catalogue scales, rather than only that something does. Scaling from 10 to 110 candidate agents or tools, the authors report routing F1 on under-specified requests dropping 16 to 23 percentage points across the models tested, and split that drop with an oracle analysis into two distinct components: a retrieval gap , the model’s failure to surface the correct tool at all, and a confusion gap , a roughly 10-percentage-point reduction in the theoretical best-case score that persists even when retrieval is assumed perfect [ 11 ] . Written as a decomposition of the error a shortlist of size m drawn from a full catalogue of n tools produces, where tool⋆\mathrm{tool}^\star is the one the request actually calls for and SmS_m is the shortlist shown to the model: error(m)  =  Pr⁡[tool⋆∉Sm]⏟retrieval gap, shrinks as m grows  +  Pr⁡[miscall∣tool⋆∈Sm]⏟confusion gap, grows as m grows\mathrm{error}(m) \;=\; \underbrace{\Pr[\mathrm{tool}^\star \notin S_m]}_{\text{retrieval gap, shrinks as } m \text{ grows}} \;+\; \underbrace{\Pr[\text{miscall} \mid \mathrm{tool}^\star \in S_m]}_{\text{confusion gap, grows as } m \text{ grows}}. Widening the shortlist drives the first term toward zero and the second term upward; narrowing it does the reverse. There is no value of m that eliminates both terms at once, which is exactly why the paper reports a design fix rather than a single ideal m : embedding-based shortlisting recovered 10 to 11 percentage points of F1 at full scale, consistently across three model families and two providers, and 10 to 17 percentage points on live production traffic, validated against 1,435 human-labeled utterances [ 11 ] . The fix is not “show fewer tools” as a blanket rule. It is choosing which few to show, per request, well.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to The Token Tax of Giving a Model More Tools

See this formula across 1 published context →

Browse the mathematical compendium →