← All parts of this equation

Equation 10 · Part 1 · Building a Multimodal AI Application That Actually Uses Its Inputs

Symbol T

T(w,h)=⌈wP⌉×⌈hP⌉,T(w, h) = \left\lceil \frac{w}{P} \right\rceil \times \left\lceil \frac{h}{P} \right\rceil,
TT

What this part means

T is part of the quantity the equation computes from the expression on the right.

Its job in the formula

T is part of the quantity the equation computes from the expression on the right.

The passage around this formula

Image and video inputs are priced by tokenization rules that are public, mechanical, and different enough across vendors that a workload’s cost cannot be estimated from a single provider’s numbers and assumed to generalise. The general form, common to every tiling or patch-based image tokenizer, is T(w,h)=⌈wP⌉×⌈hP⌉T(w, h) = \left\lceil \frac{w}{P} \right\rceil \times \left\lceil \frac{h}{P} \right\rceil. where w and h are the image’s pixel dimensions after any provider-side resize and P is that provider’s patch edge length in pixels. The formula is the same shape everywhere; the constant P , the resize rule applied before it, and the price per resulting token are what an integrator actually has to look up and budget against.

Read this part in the article →

Learn the underlying idea

A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.

Open the illustrated functions: inputs become outputs guide →

See this notation across published equations →

Sources cited in the article section

These citations provide research context; check each source for the exact claim it supports.