← Mathematical compendium

Published equation contexts

T(w,h)=⌈wP⌉×⌈hP⌉T(w, h) = \left\lceil \frac{w}{P} \right\rceil \times \left\lceil \frac{h}{P} \right\rceil

Why this formula appears here

Image and video inputs are priced by tokenization rules that are public, mechanical, and different enough across vendors that a workload’s cost cannot be estimated from a single provider’s numbers and assumed to generalise. The general form, common to every tiling or patch-based image tokenizer, is T(w,h)=⌈wP⌉×⌈hP⌉T(w, h) = \left\lceil \frac{w}{P} \right\rceil \times \left\lceil \frac{h}{P} \right\rceil. where w and h are the image’s pixel dimensions after any provider-side resize and P is that provider’s patch edge length in pixels. The formula is the same shape everywhere; the constant P , the resize rule applied before it, and the price per resulting token are what an integrator actually has to look up and budget against.

Read the full article-specific guide →

Read the representative guide

PP

Symbol P

P occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

T(w,h)=⌈wP⌉×⌈hP⌉,T(w, h) = \left\lceil \frac{w}{P} \right\rceil \times \left\lceil \frac{h}{P} \right\rceil,

Equation 10 · Foundation Models

Building a Multimodal AI Application That Actually Uses Its Inputs

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Image and video inputs are priced by tokenization rules that are public, mechanical, and different enough across vendors that a workload’s cost cannot be estimated from a single provider’s numbers and assumed to generalise. The general form, common to every tiling or patch-based image tokenizer, is T(w,h)=⌈wP⌉×⌈hP⌉T(w, h) = \left\lceil \frac{w}{P} \right\rceil \times \left\lceil \frac{h}{P} \right\rceil. where w and h are the image’s pixel dimensions after any provider-side resize and P is that provider’s patch edge length in pixels. The formula is the same shape everywhere; the constant P , the resize rule applied before it, and the price per resulting token are what an integrator actually has to look up and budget against.

Meanings in this article

  • ww: the image’s pixel dimensions after any provider-side resize.
  • hh: the image’s pixel dimensions after any provider-side resize.
Equation guide → · Article →