← All parts of this equation

Equation 5 · Part 7 · What Multimodal AI Actually Costs, Modality by Modality

subscript

costimage=Ntok(H,W)⋅ptok\mathrm{cost}_{\mathrm{image}} = N_{\mathrm{tok}}(H,W) \cdot p_{\mathrm{tok}}
subscript

What this part means

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Its job in the formula

A subscript distinguishes a version, component, step, or member of a quantity. It does not automatically mean multiplication.

The passage around this formula

with a request’s image cost simply costimage\mathrm{cost}_{\mathrm{image}} = Ntok(H,W)N_{\mathrm{tok}}(H,W) ⋅\cdot ptokp_{\mathrm{tok}} at the model’s own per-token price ptokp_{\mathrm{tok}} . The constant differs — 28 pixels for Claude, 32 for GPT-5.4, roughly 768 divided into geometry-dependent tiles for Gemini — but the shape does not: token count, and therefore cost, scales with the area of the image, not its linear size. Doubling both width and height quadruples the token bill under every one of these three schemes. That is the single fact behind essentially every dollar figure in the rest of this section, and it is also why “send a smaller image” is the one universally effective cost lever a caller has, across…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.