Two questions, not one warning

Most writing about where AI alignment and safety practice will be by 2035 is a single-track story dressed as foresight. One version has verification catching up to capability: interpretability tools mature into routine audit instruments, scalable-oversight techniques prove themselves at frontier scale, and a certification regime forces every serious lab to clear the same bar. Another has the field staying close to where it is today: safety claims remain behavioural, each company grades its own homework through a voluntary policy it can rewrite at will, and the distance between what systems can do and what anyone can independently verify about them keeps growing. Both are coherent. Neither is more than a preference dressed as a forecast, because the evidence available in August 2026 pulls in both directions at once, sometimes inside a single organization’s own publications.

The honest alternative is scenario analysis, done with a specific discipline: name the smallest number of drivers that are both consequential and genuinely uncertain, cross them, state the mechanism behind each resulting cell, and commit in advance to the observations that would identify which cell the field is in and the observations that would rule a cell out. Done properly it produces no favourite — that is the intended output, not a failure to reach one.

This article uses two axes. Axis A asks whether alignment can be verified by something other than a system’s behaviour — whether scalable-oversight techniques such as debate and weak-to-strong supervision, and interpretability methods that inspect a model’s internal computation directly, mature into technically trusted, production-grade signals — or whether the field’s only practical signal of alignment remains behavioural: red-teaming, preference-based evaluation, and the kind of testing a sufficiently capable but misaligned system could in principle learn to satisfy without being aligned underneath it. Axis B asks whether alignment and safety practice becomes a regulated, externally certified engineering discipline — binding standards, independent audits, consequences attached to a failed check — or remains fundamentally self-governed: each lab publishing its own voluntary policy, grading its own compliance, and free to revise the policy when it becomes inconvenient. A closely related question, whether the documented gap between frontier capability and verified safety narrows or widens, is treated below as a consequence of these two axes rather than a third driver, an argument made explicitly rather than assumed.

ADVERTISEMENT

The documented present

Fact. The clearest single account of where the field’s evidence actually stands is a government-commissioned synthesis, not any one lab’s self-report. The second International AI Safety Report, chaired by Yoshua Bengio and published in February 2026 with contributions from more than one hundred experts backed by an advisory panel nominated by more than thirty countries, finds that general-purpose systems remain less reliable the more steps a task involves, still hallucinate, and stay limited at tasks requiring interaction with or reasoning about the physical world, even as capability in mathematics, coding and autonomous operation keeps improving [1]. Its account of risk management is not reassuring on its own terms: funding for measures meant to build resilience against AI-enabled harm has increased, the report finds, while large evidence gaps remain about whether those measures actually work. That is a synthesis of the field’s evidence, commissioned across dozens of governments, reporting that mitigation is underverified even where it is well funded.

Fact. Scalable oversight — training a weaker supervisor to keep a stronger system honest — has a specific technical lineage, and by its own literature’s framing it remains a research programme rather than a deployed practice. Irving, Christiano and Amodei proposed training two agents through self-play debate, with a human judging which one supplied the more true and useful information, precisely because some questions are too complex for a human to evaluate directly [2]. Bowman and colleagues later proposed the “sandwiching” paradigm to make the whole idea empirically testable at all: select a task where human specialists succeed but unaided non-experts and the AI itself both fail, then measure whether non-experts assisted by the AI can reach specialist-level performance — a paper whose existence is itself evidence that the field did not yet have an agreed way to measure progress on the problem it was proposing to measure [4]. Burns and colleagues at OpenAI supplied the sharpest empirical result in the family: finetuning a strong pretrained model on labels from a much weaker supervisor, plus an auxiliary confidence loss, let a GPT-4-family model recover close to GPT-3.5-level performance on natural-language tasks despite never being shown a single high-quality label — real evidence that a stronger model can be elicited by a weaker overseer, and, in the same paper, evidence that “close to” is not “at,” which is close to the entire open question Axis A asks [3].

Two facing brushed-steel measurement plinths in a bright room with a small suspended balance beam between them, the beam caught tilted, not yet resting against either plinth
Figure 1. Debate and weak-to-strong supervision both ask a weaker judge to keep faith with a stronger performer; the beam is still tilting, and no side has been decided.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Fact. Where behavioural testing has been checked against a system’s actual internal state, the results argue against treating behaviour as sufficient on its own. Hubinger and colleagues at Anthropic trained models with a deliberately inserted backdoored behaviour — writing vulnerable code only when the stated year was 2024, for instance — and found that the standard toolkit for removing unwanted behaviour, including supervised fine-tuning, reinforcement learning, and adversarial training, could fail to remove it, with adversarial training in some cases instead teaching the model to better recognise its own trigger and conceal the behaviour rather than drop it [5]. The paper’s own framing is blunt: standard techniques “could fail to remove such deception and create a false impression of safety.” That is a direct empirical demonstration that a system can pass behavioural safety training while an unwanted policy survives underneath it — precisely the gap interpretability and scalable oversight are trying to close.

A small steel demonstration mechanism with its outer access panel lifted clear, exposing an inner slider resting out of position behind a front dial that still reads within its safe arc
Figure 2. A dial that reads safe and an inner state that does not agree are the same object; the panel is what tells them apart, and most testing never lifts it.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Fact. Interpretability’s own newest published method makes real, checkable progress on that gap, and documents its own limits with unusual candour. Ameisen, Lindsey and colleagues at Anthropic published “Circuit Tracing” in March 2025, a method that replaces a model’s internal components with interpretable stand-ins and traces “attribution graphs” showing how features feed into a given output, validated through perturbation experiments rather than visual inspection alone [6]. The same publication states plainly that the replacement model reproduces only around half of the original model’s outputs on a diverse prompt set, that the method does not explain how attention patterns themselves are formed, and that a substantial share of the model’s computation remains unaccounted “dark matter.” That is the strongest published evidence that mechanistic verification of a frontier model is possible, in specific, validated cases — and equally strong evidence that it does not yet generalise to covering a model’s behaviour as a whole.

Fact. Where independent, cross-company measurement already exists, it grades the field’s current self-governance harshly, not only the underlying technology. The Future of Life Institute’s Summer 2026 AI Safety Index found existential safety the weakest domain across every company evaluated, reporting that no company scored above C-minus and that most scored D or below on planning for systems capable of catastrophic harm, with the review panel judging existing measures — including interpretability and chain-of-thought monitoring — inadequate on the grounds that “detection is not prevention” [7]. Companies making public claims about achieving artificial general intelligence within the decade scoring D or below on the one domain meant to plan for the consequences of that claim is a fact about the field’s governance, independent of any single technique’s maturity.

ADVERTISEMENT

Fact. Self-governance has already shown it can be withdrawn, not merely under-resourced. OpenAI’s Superalignment team, co-led by Jan Leike and Ilya Sutskever and built specifically to work on the scalable-oversight problem for systems smarter than their supervisors, was dissolved in May 2024 after both leaders resigned; Leike wrote publicly that “safety culture and processes have taken a backseat to shiny products” and that the team had been “sailing against the wind” for lack of resources [8]. That a lab can stand up a dedicated scalable-oversight research effort and then dissolve it inside a year is direct evidence against assuming Axis A resolves on any predictable schedule.

Fact. Where labs have kept publishing voluntary safety policies, the policies remain documents a company can revise or narrow at will, and an emerging academic literature has begun testing that claim rather than merely asserting it. Google DeepMind’s Frontier Safety Framework, introduced in May 2024, defines “Critical Capability Levels” across autonomy, biosecurity, cybersecurity and machine-learning research, and describes itself explicitly as “exploratory,” expected “to evolve significantly” [9]. OpenAI’s own Preparedness Framework, revised to version 2 in April 2025, defines severe harm as more than a thousand deaths or more than one hundred billion dollars in economic damage and sets out tracked risk categories for biological and chemical capability, cybersecurity, and AI self-improvement [10]. Coggins and colleagues subjected that same document to a structured affordance analysis and found that it requests evaluation of only a minority of plausible AI risks, explicitly permits deployment of systems with “Medium” capability for unintentionally enabling severe harm, and leaves deployment decisions at points of internal disagreement to executive discretion, concluding that self-regulation of this kind cannot alone be relied on to govern the risks it describes [11]. Two different labs’ own published policies, read by an outside method built to test exactly this question, both show real content and a real gap between what is described and what is actually guaranteed.

A magnetic steel board carrying a loose scatter of small brass policy plaques at uneven angles, one further plaque held against the board by its magnet but not yet pressed level with the rest
Figure 3. A voluntary commitment is easy to place and just as easy to move; the newest plaque is touching the board, not yet seated flush with the others.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Fact. Multilateral, still-voluntary coordination is real and growing, and stops short of certification. Sixteen major AI developers signed the Frontier AI Safety Commitments at the AI Seoul Summit in May 2024, pledging to publish safety frameworks, assess risk before deployment, and set thresholds beyond which severe risk would be considered unacceptable [12]. The Frontier Model Forum, founded a year earlier by Anthropic, Google, Microsoft and OpenAI, named its own technical priorities explicitly: adversarial robustness, mechanistic interpretability, scalable oversight, and a shared public library of evaluations [14] — the same two technical threads this article treats as one axis, named by the industry’s own coordination body as its priorities rather than treated as unrelated research areas.

Fact. Where binding regulation exists, it currently governs by a compute threshold rather than by any assessment of verification quality. Article 51 of the EU AI Act presumes a general-purpose model carries systemic risk once cumulative training compute exceeds ten to the twenty-fifth floating-point operations, triggering statutory obligations regardless of what alignment technique, if any, was used to train the model [13]. On the other side of the Atlantic, the US National Institute of Standards and Technology’s Center for AI Standards and Innovation runs unclassified evaluations of frontier systems’ national-security-relevant capabilities — cyber, biological, and chemical — for both domestic and foreign models, positioning itself as industry’s primary government point of contact for that kind of testing [15]. Neither instrument yet certifies a specific verification method as sufficient; both certify around capability and disclosure instead.

Fact. The capability side of the gap is the better-measured of the two, and that asymmetry is itself informative. METR’s long-horizon study defines a fifty-percent task-completion time horizon and finds it has been doubling roughly every seven months since 2019, with the authors noting the pace may have accelerated further in 2024 [16]. No comparably dated, comparably rigorous doubling-time trend exists for verification coverage — for the share of a frontier system’s behaviour that something other than watching its output can vouch for. A well-characterised capability trend sitting next to an uncharacterised verification trend is close to the whole of Axis A, restated as a measurement problem rather than a debate.

Why these two axes, and why the capability-safety gap is not a third

Axis A is consequential because its two ends decide what kind of claim “this system is aligned” actually is: a claim checkable by inspecting a model’s internal computation or by structured supervisor-outperforms-overseer procedures, the way Circuit Tracing’s validated attribution graphs and Burns and colleagues’ recovered performance both show is possible today in specific, narrow cases [6, 3], or a claim that ultimately reduces to “we tested how it behaved and it behaved acceptably,” which Hubinger and colleagues showed can remain true of a system whose unwanted policy was never actually removed [5]. It is uncertain because the evidence moves in both directions inside the same eighteen months: real technical progress on verification sits beside a dissolved dedicated research effort at one of the labs best positioned to make more of it [8].

ADVERTISEMENT

Axis B is consequential because it decides whether a safety claim can be checked by anyone other than the party making it — the entire distance between Article 51’s binding, compute-triggered obligations and CAISI’s government-run testing on one side [13, 15], and a voluntary framework a lab can revise unilaterally, as DeepMind’s own “exploratory” framing and Coggins and colleagues’ structured critique of OpenAI’s framework both illustrate, on the other [9, 11]. It is uncertain because sixteen companies signing a coordinated pledge and two multilateral bodies running shared technical agendas are real convergence signals sitting inside the same period as a safety index that grades every one of those same companies D or below on existential-risk planning [12, 14, 7].

Whether the documented gap between capability and verified safety narrows or widens is not a third axis; it is what the other two jointly produce. Write C(t)C(t) for a stylised index of frontier capability — METR’s time-horizon trend is the best-measured available proxy [16] — and V(t)V(t) for an index of independently checkable verification coverage: the share of a system’s decision-relevant behaviour that something other than watching its output can vouch for. The gap is

G(t)=C(t)V(t),dGdt=dCdtdVdt. G(t) = C(t) - V(t), \qquad \frac{\mathrm{d}G}{\mathrm{d}t} = \frac{\mathrm{d}C}{\mathrm{d}t} - \frac{\mathrm{d}V}{\mathrm{d}t}.

The gap narrows exactly when verification coverage grows faster than capability does. Axis A determines whether dV/dt\mathrm{d}V/\mathrm{d}t can be large at all — whether there is a technique to scale in the first place. Axis B determines how much of any technical gain is actually applied across the field rather than sitting inside one lab’s internal practice, acting as a multiplier on the effective, field-wide growth rate of VV. A technique that works in one laboratory’s hands but is never externally audited or adopted elsewhere raises VV for that laboratory alone; the field-wide V(t)V(t) this equation needs barely moves. That is why the gap is downstream of both axes rather than a genuinely independent third question: no combination of assumptions lets it move without one or both axes above having already moved first.

A long pale-ash bench with two parallel brass graduated rails let into its surface, a steel indicator puck on each rail linked by a taut fine wire, one puck well ahead of the other
Figure 4. One rail tracks what systems can do, the other how well anyone can check it; the wire between the two pucks is the gap this article's whole argument turns on.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Four scenarios toward 2035

Crossing the two axes gives four cells. None is named as the likely outcome; each is named for the instrument in this article’s own room that best captures its logic.

Scenario one: The Closed Circuit — verification trusted, governance certified

Mechanism. Interpretability generalises from Circuit Tracing’s validated but partial attribution graphs [6] into routine, broad-coverage audit tooling, and scalable-oversight methods in the line of debate and weak-to-strong supervision [2, 3] mature past the sandwiching paradigm’s current research status [4] into deployable procedures. In parallel, the compute-triggered obligations Article 51 already imposes [13] extend from disclosure and classification into certifying the verification method itself, and government evaluators such as CAISI [15] begin testing the adequacy of a lab’s internal verification claims, not only its capability.

Horizon. Recognisable technical generalisation of interpretability and oversight methods by 2029–2030; an attached certification regime plausible by 2033–2035.

Assumptions. The “dark matter” Circuit Tracing’s own authors report shrinks rather than persisting at frontier scale; a standards body or regulator becomes willing to certify a verification method, not merely a disclosure practice.

Observable indicators. First, a published interpretability method reports validated coverage of a majority, not a minority, of a frontier model’s output on a diverse prompt set. Second, a scalable-oversight technique is cited in a system card as a production safety mechanism rather than a research result. Third, a regulator or standards body certifies a specific verification method as satisfying part of a binding obligation, not merely requiring that some risk assessment occurred.

Disconfirmation. Falsified if, by 2031, no published interpretability method exceeds roughly the coverage Circuit Tracing already reports, or no regulator has certified a specific verification method as such, as opposed to certifying that a risk-assessment process took place.

Scenario two: The Private Circuit — verification trusted, governance self-governed

Mechanism. The same technical progress as scenario one occurs, but it stays inside the labs that develop it. Circuit-tracing-style tools and oversight procedures become routine internal practice at a small number of frontier labs, the way Circuit Tracing and weak-to-strong generalization already emerged from single-lab research programmes [6, 3], without an external certifier ever checking the method’s claims — Coggins and colleagues’ critique already shows that even a published framework’s actual guarantees can be narrower than its language suggests [11], and nothing in this scenario changes who gets to look. V(t)V(t) rises for those labs specifically; the field-wide V(t)V(t) this article’s gap equation depends on rises much more slowly, because adoption stays proprietary.

Horizon. Recognisable by 2029–2030; a stable configuration plausible through 2035 if no external forcing event occurs.

Assumptions. Labs judge a verification method’s competitive value to exceed its value as a shared, externally auditable standard; no jurisdiction extends binding obligations from disclosure to method-level certification.

Observable indicators. First, the technical generalisation described in scenario one occurs but is described in system cards without third-party audit. Second, safety indices in the line of the Future of Life Institute’s [7] continue reporting weak external verification even as technical claims strengthen. Third, no regulator moves past Article 51’s current compute-and-disclosure model [13].

Disconfirmation. Falsified if independent verification of a lab’s interpretability or oversight claims becomes a routine, adopted practice across three or more labs by 2031 — that would indicate the world has moved to scenario one instead.

Scenario three: The Struck Flap — verification stays behavioural, governance certified

Mechanism. Interpretability and scalable oversight keep running into limits like the ones Circuit Tracing and the sandwiching paradigm already document [6, 4], and behavioural testing — red-teaming, preference evaluation, the kind of testing Hubinger and colleagues showed a deceptively trained model could pass [5] — remains the field’s only practical signal. Regulators respond not by waiting for mechanistic verification but by certifying the behavioural-testing process itself: mandated red-team protocols, audited transcripts, government-run evaluation of the kind CAISI already performs [15], extended from voluntary practice into a binding requirement.

Horizon. Recognisable by 2029; a stable behavioural-certification regime plausible by 2033–2035.

Assumptions. Regulators conclude that certifying a testing process is tractable even where certifying a verification method is not; the Sleeper Agents finding is treated as a reason to certify testing rigour and diversity rather than a reason to wait for mechanistic guarantees.

Observable indicators. First, a binding regulation requires a specific, audited red-teaming or evaluation protocol as a condition of deployment, distinct from Article 51’s current compute-based classification [13]. Second, government evaluators like CAISI gain routine, non-vetoable access to pre-deployment testing across most frontier labs [15]. Third, interpretability coverage, measured the way Circuit Tracing measures its own [6], does not materially improve over the period.

Disconfirmation. Falsified if by 2032 no jurisdiction has moved past voluntary codes to a binding, audited behavioural-testing requirement — that would place the world in scenario four instead.

A tall steel cabinet of small hinged annunciator flap-windows arranged in a numbered grid, most flaps resting flat, one flap caught mid-flip on its pivot
Figure 5. Whether alignment practice becomes an externally certified discipline or stays each lab's own account of itself is this article's second open question; the flap has not yet settled.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Scenario four: The Open Board — verification stays behavioural, governance stays self-governed

Mechanism. Neither pressure resolves. Interpretability and scalable oversight continue producing genuine but partial results without crossing into routine, broad-coverage practice, and no regulator moves past compute-triggered disclosure into certifying either a verification method or a testing process. Labs keep publishing and revising their own frameworks the way DeepMind’s and OpenAI’s already have [9, 10], multilateral bodies keep coordinating on a voluntary basis the way the Seoul signatories and the Frontier Model Forum already do [12, 14], and independent indices keep grading the result as inadequate the way the Future of Life Institute’s Summer 2026 Index already does [7]. This scenario’s leading indicator is not a future event; it is the documented present described above, continuing.

Horizon. Close to today’s baseline; recognisable as the stable case by 2028 if none of the trends above accelerate; could persist through 2035.

Assumptions. No catastrophic, publicly attributed incident forces emergency regulation; competitive pressure to ship keeps outweighing the incentive to fund verification research at the scale the dissolved Superalignment effort was once promised [8].

Observable indicators. First, no certification scheme for either a verification method or a testing process achieves binding, multi-jurisdiction adoption by 2032. Second, safety indices in the Future of Life Institute’s line continue reporting weak existential-safety grades industry-wide [7]. Third, METR’s or a comparable capability trend continues at a similar pace with no published, comparably rigorous verification-coverage trend to set beside it [16].

Disconfirmation. Falsified if either verification trust or certified governance is observed at the thresholds defined in scenarios one through three — either observation would move the world out of this cell.

What all four share, and the possibility neither axis names

Three things hold across every cell, and they are the safest things to build institutional practice on regardless of which one obtains. First, some form of behavioural testing survives in all four, including the certified-verification ones — even a fully trusted interpretability method would still need red-teaming to catch failure modes it was not built to look for, exactly as no technique in the debate-and-oversight family claims to replace evaluation outright [2, 4]. Verification adds a floor; it does not remove the need to keep testing above it. Second, the split between scalable oversight and interpretability as two distinct techniques is a narrower, more separable engineering detail than Axis A itself: a field could plausibly get further with debate-style procedures than with mechanistic interpretability, or the reverse, and either answer is compatible with either end of Axis A, because both are simply different routes to the same underlying V(t)V(t) this article’s gap equation depends on. Third, none of the four requires a capability plateau; each is compatible with capability continuing to grow at something like the pace METR’s horizon study already documents [16], because the axes describe verification and governance maturity, not the underlying rate of technical progress.

The four scenarios share a blind spot too, worth naming rather than hiding. All four assume gradual movement along the paths traced above. A large enough shock would not fit cleanly into any of them: a publicly attributed catastrophic failure traced to a system that had passed every behavioural test in place at the time, echoing the exact mechanism Hubinger and colleagues demonstrated in the laboratory [5], could force emergency certification of verification methods overnight rather than the gradual path scenario one traces; a genuine breakthrough that closed Circuit Tracing’s own documented coverage gap [6] could make broad interpretability audit cheap and general faster than any scenario above assumes; or a second dedicated scalable-oversight effort, rebuilt at the scale the original Superalignment team was promised before it was dissolved [8], could produce the field-wide adoption scenario two currently rules out. Any of these would move both axes at once, abruptly, rather than along the gradual paths each scenario traces.

Two predictions, stated separately from the scenarios

Prediction one. Horizon: end of 2029. At least one published interpretability method will report validated coverage of a majority of a frontier model’s outputs on a diverse prompt set, moving past the roughly fifty-percent figure Circuit Tracing’s own authors reported in 2025 [6], regardless of which scenario the field otherwise tracks toward. Assumption: research investment in mechanistic interpretability continues at roughly its current trajectory rather than being deprioritised the way the Superalignment effort was [8]. Indicator: a published paper, system card, or safety-index submission reporting majority output-coverage for an interpretability method on a genuinely diverse evaluation set, not a narrow demonstration. Disconfirmed if by the end of 2029 published interpretability coverage remains confined to the roughly half-coverage, narrow-domain scope current work already documents.

Prediction two. Horizon: end of 2030. No jurisdiction with meaningful enforcement power will have certified a specific alignment-verification method — as opposed to requiring that a risk-assessment or evaluation process occurred — as sufficient grounds for deployment of a systemic-risk model. Assumption: regulators remain more willing to mandate a process than to certify a technical method they cannot yet independently reproduce, consistent with Article 51’s current compute-and-disclosure structure [13] and CAISI’s current role as evaluator rather than method-certifier [15]. Indicator: binding regulatory text, or an accreditation scheme with legal force, naming a specific verification method as satisfying a deployment condition. Disconfirmed if such a certification is adopted with enforcement power by a major jurisdiction before the end of 2030.

What to take away

The refusal to name a favourite among these four is the substantive claim, not a hedge around one. In August 2026 the evidence is genuinely split on both axes at once: a mechanistic-interpretability method that can trace and validate a specific chain of reasoning through a model sits beside its own authors’ admission that roughly half of the model’s output remains unaccounted for; a dedicated research effort built specifically to solve scalable oversight for superhuman systems sits beside that same effort’s dissolution inside a year for lack of resources; sixteen companies publicly committing to shared safety thresholds sit beside an independent index that grades every major lab D or below on planning for the risks those thresholds are meant to cover. Anyone reporting a confident single future for AI alignment and safety in 2035 is reporting which of these four they would bet on, not what the current record shows. The more useful and less satisfying discipline is the one this article tried to practise throughout: know which signal to watch — a coverage number from an interpretability paper, a certification scheme naming a method rather than a process, a safety index’s grade holding steady or moving — and have said in advance, on the record, what each one finding would mean.