Two questions, not one storyline

Most writing about where Anthropic will be in 2035 is a single-track story wearing the costume of analysis. One version has the current safety-training method scaling smoothly to very capable systems, the Responsible Scaling Policy quietly becoming how the whole industry operates, and Claude comfortably ahead. Another has the method hitting a wall, the policy staying one company’s internal paperwork, and a rival house style winning the market instead. Both are coherent. Neither is more than a preference dressed as a forecast, because the evidence available in August 2026 does not discriminate between them.

The honest alternative is scenario analysis: name the smallest number of drivers that are both consequential and genuinely uncertain, cross them, state the mechanism behind each resulting cell, and commit in advance to the observations that would identify the cell and the observations that would kill it outright. Done properly, it produces no favourite — that is the output, not a failure to reach one.

This article uses two axes. Axis A asks where the next decade of safety training comes from: continued scaling of the constitutional and RLHF paradigm already in use, or a genuinely new method that displaces it as the primary lever. Axis B asks what the Responsible Scaling Policy architecture becomes: a shared floor, enforced by competitors adopting compatible commitments and regulators codifying comparable requirements, or a distinguishing, company-specific practice. Transparency and market position are treated here as consequences of those two axes rather than two further independent drivers — a reduction argued for explicitly below, not assumed.

ADVERTISEMENT

The documented present

Fact. Anthropic’s Responsible Scaling Policy has been rewritten three times inside twelve months. The 2023 original committed publicly not to train or deploy models capable of catastrophic harm without safety measures holding risk below acceptable levels. Version 3.0, released 24 February 2026, dropped that language: it separated unilateral commitments from what it now calls industry-wide recommendations, replaced the promised hard pause with nonbinding public goals graded through a Frontier Safety Roadmap, and introduced Risk Reports, roughly every three to six months, justifying any decision to accept marginal risk [2]. Further revisions brought the policy to version 3.4 by 8 July 2026, still built on AI Safety Level standards: ASL-2 governs current general-purpose models, and crossing named Capability Thresholds requires upgrading to the ASL-3 Security and Deployment Standards [1].

Anthropic’s stated reasoning deserves quoting rather than summarising away: it wrote that capability thresholds had proven “far more ambiguous than we anticipated,” that government action had moved slower than expected amid a shift “toward prioritizing AI competitiveness,” and that higher safety tiers might be “outright impossible to implement without collective action” [2] — a vendor’s account of its own decision, not an independent finding. An analysis from the Centre for the Governance of AI credits the Risk Reports as meaningful progress, but flags that Anthropic’s own RAND Security Level 4 protection against state-level theft of model weights moved from mandatory to merely recommended, and warns weaker competitors could lower their commitments without matching the new transparency measures [3]. The episode is the clearest evidence of how uncertain Axis B is: the policy’s architect just demonstrated its own commitments are revisable.

The threshold-triggered structure underneath all three versions is stable even as what triggers what has moved. Write R(m)R(m) for a model’s assessed risk on a given threat class and τ3,τ4\tau_3, \tau_4 for the thresholds separating safety tiers:

Safeguard(m)={ASL-2,R(m)<τ3ASL-3,τ3R(m)<τ4ASL-4,R(m)τ4 \text{Safeguard}(m) = \begin{cases} \text{ASL-2}, & R(m) < \tau_3 \\ \text{ASL-3}, & \tau_3 \le R(m) < \tau_4 \\ \text{ASL-4}, & R(m) \ge \tau_4 \end{cases}

What changed between versions is not this shape but the codomain attached to crossing a threshold. Under the original policy, crossing the top implied a commitment to halt. From version 3.0 onward, the same crossing implies an obligation to publish a Risk Report and justify the decision publicly — a disclosure requirement, not a hard constraint. That distinction is the entire disagreement about whether the architecture is strengthening or weakening. Axis B is not whether Anthropic “has” a scaling policy — it already does, on every version — but what obligation attaches to the top of the function, and whether that obligation is Anthropic’s alone or shared.

Current deployed models sit at ASL-3. Claude Opus 5, released 24 July 2026, carries chemical and biological uplift capability in the first tier Anthropic tracks but not the second, the same ASL-3 protections as its predecessor, no new concerning alignment properties in Anthropic’s own evaluation, and no crossing of the policy’s automated AI research-and-development threshold [12]. Claude Sonnet 5, released 30 June 2026, launched with cyber safeguards on by default, deliberately kept short of strong offensive cyber capability, and is reported as an improvement on its predecessor in prompt-injection resistance and hallucination rate [13]. These are vendor-reported outcomes, not independently replicated findings, a qualification that applies every time they recur below.

ADVERTISEMENT
An open ring binder of printed principles with one page caught mid-turn on a light-oak desk, beside a small interpretability probe rig whose fine debug clip hovers just short of a chip die pad
Figure 1. The current paradigm is a written set of principles the model critiques itself against; the candidate replacement is a clip that has not yet landed on the die.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Fact. The training method under these models has a public description that has not been publicly superseded. Constitutional AI trains a model to critique and revise its own outputs against written principles, then uses reinforcement learning from AI-generated rather than exclusively human-generated feedback, aiming for an assistant that explains its objections rather than refusing silently [4]. It sits alongside RLHF proper: preference modelling plus iterated online training against weekly human judgments, with Anthropic’s own published research reporting a roughly linear relation between reward and the square root of the KL divergence from the model’s starting point, one measure of how far fine-tuning can push a model before it stops resembling its base [5]. No subsequent system card has announced a different foundational method; what has changed publicly is scale and evaluation coverage, not the paradigm’s name.

Anthropic’s own 2023 statement of its safety philosophy is candid about that method’s limits and deserves quoting rather than paraphrase into false confidence. It said it does not “yet have a solid understanding of how to ensure that… powerful systems are robustly aligned with human values,” described optimistic, intermediate, and pessimistic scenarios for how hard alignment might prove, and said current techniques including RLHF and Constitutional AI might be “largely sufficient” only in the optimistic case [7] — Anthropic assessing its own paradigm’s ceiling, unretracted three years later, and the strongest evidence that Axis A is genuinely open.

Fact. Interpretability research is a second, independent line of evidence, pointing the same way. Anthropic’s interpretability team trained sparse autoencoders with up to 34 million learned features on Claude 3 Sonnet’s middle-layer activations, extracting features that generalise across languages and across text and images, for concepts from concrete entities to abstractions like sarcasm and deception, and showed that clamping one such feature reliably steers output [6]. The same publication reports that even at 34 million features, roughly two-thirds were “dead” after training and the dictionary did not capture everything the model represents. Interpretability in 2026 can reliably find and manipulate individual concepts inside a production model; it cannot yet offer anything close to a complete account of what the model is doing, and the paper says so.

Fact. Convergence infrastructure for Axis B already exists in two forms, neither yet a shared binding standard. Anthropic, Google, Microsoft, and OpenAI founded the Frontier Model Forum on 26 July 2023 with stated goals including coordinated research on interpretability and scalable oversight and a public library of evaluations — goals for future coordination, not commitments already unified [9]. Separately, the EU’s General-Purpose AI Code of Practice, effective 2 August 2025, requires any provider presumed to carry systemic risk — above ten-to-the-twenty-fifth floating-point operations of training compute, roughly eleven providers worldwide — to maintain a state-of-the-art safety framework before release, run structured risk identification, and publish a mandatory safety report [11]. That architecture sits structurally close to Anthropic’s own, though the Code never names the Responsible Scaling Policy — suggestive of convergence, not proof of it.

A shelf of identically bound safety-framework binders with one differently bound volume held out at the gap where it would sit, its spine turned away and not yet slotted in
Figure 2. A framework becomes a floor when it looks the same on every shelf; this one is still being held at the gap, not yet set down.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Fact. A third, newer form of convergence infrastructure is external, non-vendor-controlled evaluation. METR’s Frontier Risk Report, window 16 February to 16 March 2026, published 19 May 2026, had Anthropic, Google, Meta, and OpenAI grant access to raw chain-of-thought reasoning, non-public capability information, and internal control-protocol details, to assess whether internal AI use at any participant could produce an unauthorised “rogue deployment.” Participants controlled which proprietary details appeared in public but could not veto the conclusions [10] — editorial independence that is itself a leading indicator, since a practice surviving repeated rounds without vendor sign-off differs meaningfully from a one-off audit granted for appearances.

Fact. Anthropic’s market position has moved sharply within this article’s evidence baseline. The company raised sixty-five billion dollars in a Series H round announced 28 May 2026 at a nine-hundred-sixty-five-billion-dollar valuation, with run-rate revenue crossing forty-seven billion dollars that month — up from roughly one billion annualised in December 2024 [8]. Independent measures agree on direction: Menlo Ventures’ mid-2025 update put Anthropic at thirty-two percent of enterprise large-language-model usage against OpenAI’s twenty-five, reversing OpenAI’s fifty percent at the end of 2023 [14]; Ramp’s payments-based index recorded Anthropic passing OpenAI for the first time in April 2026, thirty-four-point-four percent to thirty-two-point-three [15]. A funding disclosure, a usage survey, and a transaction index agreeing on direction is the right amount of confidence to place in any one of them alone.

ADVERTISEMENT

Why two axes, and why transparency and market position are not a third and fourth

Analysis. An axis earns inclusion by being both consequential and genuinely uncertain. Axis A is consequential because the two ends imply different institutional realities: continued scaling of Constitutional AI and RLHF-style post-training keeps safety work inside the same evaluation-and-reinforcement loop already legible to the people running it, while a genuinely new paradigm means red-teaming practice, evaluation suites, and the whole Responsible Scaling Policy threshold apparatus — calibrated to the current method — would need substantial rebuilding. Axis A is uncertain because Anthropic’s own 2023 assessment left the question open in three directions [7], and interpretability evidence three years later shows real capability without closing the gap that assessment described [6].

Axis B is consequential because it decides whether the threshold-and-report architecture becomes something a regulator or competitor can be held to, or stays one company’s published internal document. It is uncertain because the evidence points opposite ways inside the same twelve months: Anthropic loosened its hardest commitment in the same period the European Union hardened a structurally similar one into binding law, and an independent evaluator ran its first no-veto assessment across four companies in that same window [2, 11, 10].

The axes are not fully independent. A rupture on Axis A would likely force a rewrite of Axis B’s threshold apparatus regardless of cell, since thresholds calibrated to the current method’s failure modes would not transfer cleanly. Causation runs the other way too: convergence arriving first through binding regulation, rather than voluntary coordination, could lock in evaluation practice built around the current paradigm and make a later rupture harder to comply with, not easier. The axes constrain without determining each other — the condition under which crossing them is informative rather than decorative.

Transparency and market position are treated as downstream, because neither has an uncertain-and-consequential story of its own that is not already a restatement of Axis A or B. Whether Anthropic discloses more or less is really whether disclosure is externally compelled (Axis B) and how much is disclosable given how well the paradigm is understood from the inside (Axis A) — Scaling Monosemanticity shows transparency about mechanism is bottlenecked on interpretability research, Axis A’s territory, not a separate lever [6]. Whether Claude’s market position leads or lags is a signature each scenario predicts rather than a third axis, because nothing above ties revenue or enterprise share mechanically to either training method or governance architecture — a company can plausibly win commercially in any of the four cells below.

Four scenarios toward 2035

Scenario one: Common Constitution — paradigm continues, policy converges

Mechanism. Constitutional AI and RLHF-style post-training keep working well enough that no rival method displaces them as the primary safety lever. Simultaneously, the Responsible Scaling Policy’s threshold-and-report architecture becomes a genuine cross-industry floor — not necessarily through voluntary matching, but because instruments like the EU Code of Practice harden further, other jurisdictions adopt comparable thresholds, and evaluators like METR run recurring, non-vetoable assessments across every major lab as a condition of enterprise and government access. Anthropic’s governance choices stop differentiating it commercially because everyone serious clears a similar bar.

Horizon. Full convergence recognisable by 2033–2035; the leading indicators below should be visible by 2029–2030 if this scenario is underway.

Assumptions. No paradigm-breaking method reaches production at any major lab; at least one more jurisdiction with enforcement power extends systemic-risk obligations comparable to the EU Code of Practice; non-vetoable external evaluation becomes repeated rather than one-off.

Observable indicators. First, two or more frontier labs besides Anthropic publish Risk-Report-shaped, periodic, threshold-triggered documents before 2030. Second, a jurisdiction outside the EU adopts a compute- or capability-based systemic-risk threshold with mandatory pre-release reporting, not just voluntary guidance, before 2030. Third, METR or a comparable evaluator runs a third or fourth round of non-vetoable assessment across four or more companies with no participant suppressing an unfavourable conclusion.

Disconfirmation. This scenario is dead if, by 2031, no frontier lab besides Anthropic has published a comparably structured public risk report, and no jurisdiction with meaningful enforcement power has adopted a binding systemic-risk reporting requirement — that combination would mean the convergence infrastructure documented above stalled rather than hardened.

Scenario two: House Constitution — paradigm continues, policy stays Anthropic’s own

Mechanism. The same technical continuation holds, but the governance architecture around it does not generalise. Regulatory codification stalls or fragments across incompatible jurisdictions; competitors do not adopt comparably structured frameworks, or adopt ones that diverge enough that no external party can hold two companies to the same bar. The Responsible Scaling Policy becomes what version 3.0’s own language already gestures toward: a recommendation the industry does not take up, alongside a smaller set of things Anthropic keeps committing to unilaterally [2]. Governance becomes a brand attribute rather than a shared floor — something to point to competitively, not something a regulator or rival can be measured against.

Horizon. Recognisable divergence by 2030; a stable, non-converging steady state plausible by 2035.

Assumptions. The current paradigm keeps scaling without a public rupture; government appetite for binding cross-industry regulation weakens or stays fragmented, consistent with the “prioritizing AI competitiveness” shift Anthropic itself cited [2]; competitors treat safety commitments as market positioning rather than a compliance floor.

Observable indicators. First, the count of labs publishing Risk-Report-equivalent disclosures stays at one — Anthropic — through 2030. Second, the RAND Security Level 4 downgrade stays unreversed and unmatched by any competitor adopting an equivalent standard as mandatory. Third, systemic-risk regulation stays confined to the EU Code of Practice, with no comparable binding requirement in the United States, United Kingdom, or major Asian markets.

Disconfirmation. Falsified if, before 2031, two or more competitors adopt Risk-Report-equivalent public disclosure as a matter of routine practice rather than one-off response to incident, or a second major jurisdiction adopts binding systemic-risk reporting requirements comparable in force to the EU Code of Practice — either would indicate the floor is generalising rather than staying Anthropic’s alone.

A bound capacity-planning printout open on a desk in front of a frosted-glass partition, beyond which one server unit in a small rack row sits with its blanking plate off and its status strip still dark while its neighbours glow softly
Figure 3. Every scenario still has to be paid for and powered; the unit still dark behind the glass is capacity ordered but not yet drawing load.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Scenario three: New Compact — paradigm ruptures, policy converges

Mechanism. Something displaces Constitutional AI and RLHF-style post-training as the primary lever — plausibly a method built natively around interpretability rather than behavioural preference, or a scalable-oversight method such as AI-assisted debate that the Frontier Model Forum’s own research agenda already names as a priority [9]. Because the rupture makes every lab’s existing threshold calibration obsolete at once, it becomes the forcing event for standardisation: regulators and competitors converge because of the rupture, not despite it, since no single company’s evaluation infrastructure survives the transition intact and a shared rebuild is cheaper than several incompatible ones.

Horizon. The rupture itself plausible any time from 2028 onward given how open Axis A already is; full convergence around its replacement plausible by 2035 if the rupture happens early enough in the window to allow rebuilding shared infrastructure.

Assumptions. Interpretability crosses from feature-level explanation — the 34-million-feature dictionaries already demonstrated, most features still dead or uncatalogued [6] — to a direct training signal rather than only a post hoc audit tool; the rupture is legible and public enough for competitors and regulators to coordinate around rather than each absorbing it privately.

Observable indicators. First, a system card from any major lab describes a training method whose primary safety mechanism is not describable as constitutional self-critique or preference-based reinforcement learning. Second, interpretability research demonstrates a method that measurably reduces the “dead feature” problem Scaling Monosemanticity reported, rather than just training larger dictionaries with the same limitation. Third, the Responsible Scaling Policy or a successor is restructured around a materially different threshold-setting method within eighteen months of such a disclosure, and at least one competitor adopts a similar restructuring in the same window.

Disconfirmation. Falsified if, by 2032, every major lab’s published system cards still describe training methods recognisably descended from constitutional self-critique and RLHF, regardless of scale or effort-control sophistication — the absence of a rupture is itself sufficient to place the world in scenario one or two instead.

A steel straightedge holding down a printed evaluation transcript on a bench, a highlighter caught mid-stroke with its cap off and its ink trail stopping partway across the line
Figure 4. A safety claim is only as good as the outside check on it; the stroke stopping mid-line is the check still being made, not yet signed off.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Scenario four: Proprietary Break — paradigm ruptures, policy stays company-specific

Mechanism. A genuinely new safety-training method emerges as a competitive moat rather than a shared standard. The lab that develops it treats the method, not just its outputs, as proprietary, reasoning that a safety-training breakthrough is also a capability and trust advantage worth defending commercially. Other labs cannot replicate it, replicate an inferior version, or build incompatible alternatives. Regulators, lacking the technical basis to write requirements around a method only one or two companies understand, regulate outputs and incidents instead, or fall behind. Fragmentation deepens rather than resolves — the rupture that could have forced convergence instead makes it impossible, and evaluation suites, threshold definitions, and safety claims across labs stop being comparable at all, not merely differently structured.

Horizon. A rupture on this path plausible from 2028 onward, same as scenario three; the divergence it produces widening through 2035 rather than closing.

Assumptions. Whichever lab develops the new method has commercial incentive and legal ability to keep it undisclosed beyond regulatory minimums; a market already this fast-moving — Anthropic’s own reported forty-sevenfold revenue growth in seventeen months is one data point [8] — rewards guarding a breakthrough rather than sharing it; regulatory capacity to write technically literate requirements around an undisclosed method keeps lagging disclosed capability, consistent with how narrowly the EU Code of Practice’s Safety chapter is scoped to compute thresholds rather than method [11].

Observable indicators. First, a system card discloses a materially different safety outcome — different alignment-evaluation results, different refusal or deception characteristics — without disclosing a comparably different method, and competitors cannot reproduce it. Second, cross-lab safety comparisons previously possible on a shared basis become explicitly non-comparable in independent evaluators’ own reporting. Third, regulatory language shifts toward outcome- and incident-based requirements rather than method-based ones, an implicit admission that method-level requirements are no longer enforceable.

Disconfirmation. Falsified if a paradigm rupture occurs but the method or a functionally equivalent one is independently reproduced, disclosed, or standardised by a second lab or a standards body within three years of first appearing — rapid replication would place the world in scenario three instead, regardless of the original developer’s intentions.

What all four share, and the possibility neither axis names

Analysis. Three things hold across every cell, and they are the safest things to build institutional practice on regardless of which scenario obtains. First, some version of the threshold-triggered safeguard function survives in all four — the shape shown above, not necessarily its current thresholds or codomain. Even Proprietary Break, the most fragmented cell, still implies each lab running some internal version of “more assessed risk requires more safeguard,” since nothing removes the underlying engineering need; what varies is only who can see the function and who can compel a given output from it. Second, some disclosure artifact resembling a Risk Report persists in every cell, because the practice is established enough — three successive policy versions have kept it — that abandoning it outright would itself be a costly, legible signal. Third, market position is not determined by either axis alone in any of the four scenarios, which is itself the useful finding about Claude’s competitive standing: the scenarios were built so commercial success is compatible with any combination of continuity or rupture and convergence or divergence, because nothing above ties market position mechanically to either.

The far edge of the scenario-planning wall, where a single fresh kraft card is pinned alone beyond the four linked clusters with no string yet running to it, a mostly-full spool of cotton string resting nearby
Figure 5. Nothing here rules out a fifth thing neither axis names; a card off the network is what that would look like before anyone has drawn a line to it.Image prompt and art direction by Brecht Corbeel; image generated to that direction.

The four scenarios share a blind spot too, worth naming rather than hiding. All four assume the two axes here remain the two that matter. A sufficiently large exogenous shock — state nationalisation of frontier compute, a catastrophic incident triggering emergency international regulation outside the voluntary-and-EU-led track already visible, or a security failure at the scale RAND Security Level 4 was designed against and that Anthropic just downgraded from mandatory to recommended [3] — would not fit cleanly into any of the four cells, because it would move both axes at once, abruptly, and by force rather than the gradual processes each scenario assumes.

Two predictions, stated separately from the scenarios

Prediction one. Horizon: end of 2029. External, non-vetoable evaluation of internal AI use — the practice METR piloted across four companies in early 2026 — will have run at least twice more, across at least those four companies, regardless of which scenario the industry otherwise tracks toward. Assumption: no single catastrophic evaluation failure discredits the practice in its first few rounds. Indicator: publicly referenced, dated successor reports, or a functionally equivalent recurring assessment, in company or evaluator communications. Disconfirmed if by end of 2029 no comparable assessment has recurred beyond the 2026 pilot.

Prediction two. Horizon: end of 2030. The Responsible Scaling Policy will undergo at least one further material rewrite, not a minor version bump, regardless of scenario — every version to date has been revised within roughly a year of the last, and nothing above suggests that pace has structural reasons to slow. Assumption: Anthropic remains independent, still setting its own policy rather than one absorbed into a superseding external regime. Indicator: a version change with structural rather than cosmetic revision — a changed threshold-setting method, disclosure obligation, or unilateral-versus-industry split. Disconfirmed if the policy at end of 2030 is structurally the same document as version 3.4, with only threshold numbers updated.

What to take away

The refusal to pick a favourite among these four is the substantive claim, not a hedge around one. In August 2026 the evidence is genuinely split on both axes at once: a company that just removed its own hardest safety commitment in the same year a major regulator hardened a structurally similar one into binding law; a training paradigm whose own architects said three years ago they did not know if it would scale, tested since then by interpretability research that keeps finding more inside these models without finding all of it; a market position moving fast enough that commercial success looks achievable under any combination of the other two. Anyone reporting a confident single future for Anthropic in 2035 is reporting which of these four they would bet on, not what the record currently shows. The useful discipline is narrower and less satisfying: know which signal to watch, and have said in advance, on the record, what each one would mean.