Equation 1 · Ten Failure Modes That Define Small and On-Device AI Deployments
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol S_request
equest is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol i
i is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
A device can report gigabytes of free memory and still kill the process that asks for a much smaller allocation, because the operating system’s out-of-memory killer does not care about total free memory — it cares about whether a single contiguous block large enough for the request exists. Large weight tensors and growing key-value caches on a constrained device fragment memory into scattered free regions faster than they fill it, and an allocation can fail against a device that is, in aggregate, nowhere near full. is the condition that actually governs whether an allocation succeeds, and it can fail even when holds comfortably — the requested…
Read the full surrounding passage
A device can report gigabytes of free memory and still kill the process that asks for a much smaller allocation, because the operating system’s out-of-memory killer does not care about total free memory — it cares about whether a single contiguous block large enough for the request exists. Large weight tensors and growing key-value caches on a constrained device fragment memory into scattered free regions faster than they fill it, and an allocation can fail against a device that is, in aggregate, nowhere near full. is the condition that actually governs whether an allocation succeeds, and it can fail even when holds comfortably — the requested size fits inside total free memory but not inside any one free block. This is precisely the problem the smartphone inference system PowerInfer-2 was built to manage: it introduces a segmented neuron cache with a least-recently-used eviction policy specifically to keep the working set of “hot” and “cold” neuron clusters from fragmenting a mobile device’s limited DRAM, allowing models “whose sizes exceed the device’s memory capacity” to run without the allocator failures that naive loading produces, while cutting memory usage by roughly 40% relative to prior mobile inference engines at comparable speed [ 3 ] . The fix is not more free memory; it is memory management that treats fragmentation as the binding constraint rather than total capacity.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to Ten Failure Modes That Define Small and On-Device AI Deployments