Equation 5 · What Multimodal AI Actually Costs, Modality by Modality
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol N_tok
ok is one factor in the product that computes the quantity on the left.
Symbol H
H is one factor in the product that computes the quantity on the left.
Symbol W
W is one factor in the product that computes the quantity on the left.
Symbol p_tok
ok is one factor in the product that computes the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
with a request’s image cost simply = at the model’s own per-token price . The constant differs — 28 pixels for Claude, 32 for GPT-5.4, roughly 768 divided into geometry-dependent tiles for Gemini — but the shape does not: token count, and therefore cost, scales with the area of the image, not its linear size. Doubling both width and height quadruples the token bill under every one of these three schemes. That is the single fact behind essentially every dollar figure in the rest of this section, and it is also why “send a smaller image” is the one universally effective cost lever a caller has, across…
Read the full surrounding passage
with a request’s image cost simply = at the model’s own per-token price . The constant differs — 28 pixels for Claude, 32 for GPT-5.4, roughly 768 divided into geometry-dependent tiles for Gemini — but the shape does not: token count, and therefore cost, scales with the area of the image, not its linear size. Doubling both width and height quadruples the token bill under every one of these three schemes. That is the single fact behind essentially every dollar figure in the rest of this section, and it is also why “send a smaller image” is the one universally effective cost lever a caller has, across all three vendors, without changing anything about the model itself.
Sources cited in the article section
These citations give research context. Read each source to check which claims it supports.
Return to What Multimodal AI Actually Costs, Modality by Modality