← All parts of this equation

Equation 4 · Part 7 · What Multimodal AI Actually Costs, Modality by Modality

multiplication

Ntok(H,W)=⌈HP⌉×⌈WP⌉N_{\mathrm{tok}}(H,W) = \left\lceil \frac{H}{P} \right\rceil \times \left\lceil \frac{W}{P} \right\rceil
multiplication

What this part means

Multiply the quantities on either side.

Its job in the formula

Multiply the quantities on either side.

The passage around this formula

A single formula covers all three counting schemes if the tile or patch edge length is left as a free constant P , specific to the vendor: Ntok(H,W)=⌈HP⌉×⌈WP⌉N_{\mathrm{tok}}(H,W) = \left\lceil \frac{H}{P} \right\rceil \times \left\lceil \frac{W}{P} \right\rceil. with a request’s image cost simply costimage\mathrm{cost}_{\mathrm{image}} = Ntok(H,W)N_{\mathrm{tok}}(H,W) ⋅\cdot ptokp_{\mathrm{tok}} at the model’s own per-token price ptokp_{\mathrm{tok}} . The constant differs — 28 pixels for Claude, 32 for GPT-5.4, roughly 768 divided into geometry-dependent tiles for Gemini — but the shape does not: token count, and therefore cost, scales with the area of the image, not its linear size. Doubling both width and height quadruples the token bill under every one of these three schemes. That is the single fact behind…

Read this part in the article →

Learn the underlying idea

Multiplication scales one quantity by another. A dot, a cross, or adjacent symbols can indicate a product.

Open the illustrated multiplication: combining factors guide →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.