Equation 12 · Model Systems in 2035: Four Scenarios, Their Signals, and What Would Falsify Them
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol I^*
is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
superscript
A raised number can be a power. When it is a label or bound, it selects a case or the upper limit of a sum; the formula’s structure distinguishes these uses.
See an illustrated explanation →How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
They are only partly independent, and the coupling must be stated. Three linkages matter. First, the two ends of Axis A load the machine differently: large training runs are dense matrix-multiplication workloads with high arithmetic intensity that sit comfortably on the compute-bound side of , whereas long deliberation is autoregressive decoding, which is bandwidth-bound. A world that moves toward deliberation therefore makes Axis B bind harder for the same total FLOP , which is why the axes cannot be treated as orthogonal. Second, causation runs backwards too: if bandwidth relief arrives, deliberation becomes cheaper, which pushes Axis A toward the inference end. Third, both ends of…
Read the full surrounding passage
They are only partly independent, and the coupling must be stated. Three linkages matter. First, the two ends of Axis A load the machine differently: large training runs are dense matrix-multiplication workloads with high arithmetic intensity that sit comfortably on the compute-bound side of , whereas long deliberation is autoregressive decoding, which is bandwidth-bound. A world that moves toward deliberation therefore makes Axis B bind harder for the same total FLOP , which is why the axes cannot be treated as orthogonal. Second, causation runs backwards too: if bandwidth relief arrives, deliberation becomes cheaper, which pushes Axis A toward the inference end. Third, both ends of Axis A ultimately consume grid power, so the energy component of Axis B is common to all four cells and differs only in where the load sits.
Sources cited in the article section
- [7] Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters ↗
- [16] Energy and AI: Executive Summary ↗
These citations give research context. Read each source to check which claims it supports.
Return to Model Systems in 2035: Four Scenarios, Their Signals, and What Would Falsify Them