← All parts of this equation

Equation 5 · Part 2 · What Multimodal AI Actually Costs, Modality by Modality

Symbol H

costimage=Ntok(H,W)⋅ptok\mathrm{cost}_{\mathrm{image}} = N_{\mathrm{tok}}(H,W) \cdot p_{\mathrm{tok}}
HH

What this part means

H is one factor in the product that computes the quantity on the left.

Its job in the formula

H is one factor in the product that computes the quantity on the left.

The passage around this formula

with a request’s image cost simply costimage\mathrm{cost}_{\mathrm{image}} = Ntok(H,W)N_{\mathrm{tok}}(H,W) ⋅\cdot ptokp_{\mathrm{tok}} at the model’s own per-token price ptokp_{\mathrm{tok}} . The constant differs — 28 pixels for Claude, 32 for GPT-5.4, roughly 768 divided into geometry-dependent tiles for Gemini — but the shape does not: token count, and therefore cost, scales with the area of the image, not its linear size. Doubling both width and height quadruples the token bill under every one of these three schemes. That is the single fact behind essentially every dollar figure in the rest of this section, and it is also why “send a smaller image” is the one universally effective cost lever a caller has, across…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.