Equation 10 · AI Alignment and Safety in 2035: Two Axes, Four Scenarios, and What Would Falsify Them
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol V
V is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Symbol t
t is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Three things hold across every cell, and they are the safest things to build institutional practice on regardless of which one obtains. First, some form of behavioural testing survives in all four, including the certified-verification ones — even a fully trusted interpretability method would still need red-teaming to catch failure modes it was not built to look for, exactly as no technique in the debate-and-oversight family claims to replace evaluation outright [ 2 , 4 ] . Verification adds a floor; it does not remove the need to keep testing above it. Second, the split between scalable oversight and interpretability as two distinct techniques is a narrower, more separable engineering detail…
Read the full surrounding passage
Three things hold across every cell, and they are the safest things to build institutional practice on regardless of which one obtains. First, some form of behavioural testing survives in all four, including the certified-verification ones — even a fully trusted interpretability method would still need red-teaming to catch failure modes it was not built to look for, exactly as no technique in the debate-and-oversight family claims to replace evaluation outright [ 2 , 4 ] . Verification adds a floor; it does not remove the need to keep testing above it. Second, the split between scalable oversight and interpretability as two distinct techniques is a narrower, more separable engineering detail than Axis A itself: a field could plausibly get further with debate-style procedures than with mechanistic interpretability, or the reverse, and either answer is compatible with either end of Axis A, because both are simply different routes to the same underlying V(t) this article’s gap equation depends on. Third, none of the four requires a capability plateau; each is compatible with capability continuing to grow at something like the pace METR’s horizon study already documents [ 16 ] , because the axes describe verification and governance maturity, not the underlying rate of technical progress.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to AI Alignment and Safety in 2035: Two Axes, Four Scenarios, and What Would Falsify Them