← Back to article

Equation 1 · Ten Failure Modes That Define Small and On-Device AI Deployments

What does this equation mean?

Srequest≤max⁡i(freei)S_{\mathrm{request}} \le \max_i(\mathrm{free}_i)

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

SrequestS_{\mathrm{request}}

Symbol S_request

SrS_request is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

ii

Symbol i

i is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

A device can report gigabytes of free memory and still kill the process that asks for a much smaller allocation, because the operating system’s out-of-memory killer does not care about total free memory — it cares about whether a single contiguous block large enough for the request exists. Large weight tensors and growing key-value caches on a constrained device fragment memory into scattered free regions faster than they fill it, and an allocation can fail against a device that is, in aggregate, nowhere near full. is the condition that actually governs whether an allocation succeeds, and it can fail even when SrequestS_{\mathrm{request}} ≤\le ∑i\sum_i freei\mathrm{free}_i holds comfortably — the requested…
Read the full surrounding passage
A device can report gigabytes of free memory and still kill the process that asks for a much smaller allocation, because the operating system’s out-of-memory killer does not care about total free memory — it cares about whether a single contiguous block large enough for the request exists. Large weight tensors and growing key-value caches on a constrained device fragment memory into scattered free regions faster than they fill it, and an allocation can fail against a device that is, in aggregate, nowhere near full. is the condition that actually governs whether an allocation succeeds, and it can fail even when SrequestS_{\mathrm{request}} ≤\le ∑i\sum_i freei\mathrm{free}_i holds comfortably — the requested size fits inside total free memory but not inside any one free block. This is precisely the problem the smartphone inference system PowerInfer-2 was built to manage: it introduces a segmented neuron cache with a least-recently-used eviction policy specifically to keep the working set of “hot” and “cold” neuron clusters from fragmenting a mobile device’s limited DRAM, allowing models “whose sizes exceed the device’s memory capacity” to run without the allocator failures that naive loading produces, while cutting memory usage by roughly 40% relative to prior mobile inference engines at comparable speed [ 3 ] . The fix is not more free memory; it is memory management that treats fragmentation as the binding constraint rather than total capacity.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Ten Failure Modes That Define Small and On-Device AI Deployments

See this formula across 1 published context →

Browse the mathematical compendium →