Equation 4 · What Multimodal AI Actually Costs, Modality by Modality
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol N_tok
ok is part of the quantity the equation computes from the expression on the right.
Symbol H
H occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Symbol W
W occurs above the fraction bar. The numerator is divided by the entire denominator below it.
Symbol P
P occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
A single formula covers all three counting schemes if the tile or patch edge length is left as a free constant P , specific to the vendor: . with a request’s image cost simply = at the model’s own per-token price . The constant differs — 28 pixels for Claude, 32 for GPT-5.4, roughly 768 divided into geometry-dependent tiles for Gemini — but the shape does not: token count, and therefore cost, scales with the area of the image, not its linear size. Doubling both width and height quadruples the token bill under every one of these three schemes. That is the single fact behind…
Read the full surrounding passage
A single formula covers all three counting schemes if the tile or patch edge length is left as a free constant P , specific to the vendor: . with a request’s image cost simply = at the model’s own per-token price . The constant differs — 28 pixels for Claude, 32 for GPT-5.4, roughly 768 divided into geometry-dependent tiles for Gemini — but the shape does not: token count, and therefore cost, scales with the area of the image, not its linear size. Doubling both width and height quadruples the token bill under every one of these three schemes. That is the single fact behind essentially every dollar figure in the rest of this section, and it is also why “send a smaller image” is the one universally effective cost lever a caller has, across all three vendors, without changing anything about the model itself.
Sources cited in the article section
These citations give research context. Read each source to check which claims it supports.
Return to What Multimodal AI Actually Costs, Modality by Modality