← All parts of this equation

Equation 10 · Part 4 · Building a Multimodal AI Application That Actually Uses Its Inputs

Symbol P

T(w,h)=⌈wP⌉×⌈hP⌉,T(w, h) = \left\lceil \frac{w}{P} \right\rceil \times \left\lceil \frac{h}{P} \right\rceil,
PP

What this part means

P occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Its job in the formula

P occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

The passage around this formula

Image and video inputs are priced by tokenization rules that are public, mechanical, and different enough across vendors that a workload’s cost cannot be estimated from a single provider’s numbers and assumed to generalise. The general form, common to every tiling or patch-based image tokenizer, is T(w,h)=⌈wP⌉×⌈hP⌉T(w, h) = \left\lceil \frac{w}{P} \right\rceil \times \left\lceil \frac{h}{P} \right\rceil. where w and h are the image’s pixel dimensions after any provider-side resize and P is that provider’s patch edge length in pixels. The formula is the same shape everywhere; the constant P , the resize rule applied before it, and the price per resulting token are what an integrator actually has to look up and budget against.

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.