A rule has to name something
Regulating a technology begins with an act of naming. Before any requirement can be enforced, the rule must identify three things: the object the obligation attaches to, the person it binds, and the moment at which compliance is judged. For most regulated artefacts these choices are easy, because the artefact holds still. A pressure vessel is a pressure vessel, its manufacturer is stamped on the plate, and it either satisfied the standard when it left the works or it did not.
A machine-learning system supplies none of that stability. The weights are one object; the system wrapped around them is another; the service that answers a request is a third, and it can be swapped for a different one overnight while the name on the invoice stays the same. The interesting question is therefore not whether these systems should be regulated. It is a narrower and more technical one: what can a rule actually get hold of?
Five attachment points recur across every framework now in force or in draft. A rule can attach to the model as an artefact; to the role a party plays, such as provider or deployer; to the use case, defined by the consequence of the output; to the compute spent on training; or to the service as offered to the public. Each is administrable in a different degree, and each fails in a different direction. The rest of this article works through those failures, because they are where the live disputes are.
What the frameworks actually attach to
The European Union attaches to roles first, then to tiers of use, then separately to models. The AI Act defines a provider as a person or body that develops a system or a general-purpose AI model, or has one developed, and places it on the market under its own name or trademark; a deployer is one using a system under its authority outside a personal non-professional activity [1]. Obligations then attach according to the risk tier of the use: a set of prohibited practices, a high-risk category covering uses such as hiring, education access and law enforcement, a transparency tier for systems people interact with, and a residual minimal-risk class carrying no specific requirements [7]. On top of that sits a distinct layer keyed to the model rather than the use: providers of general-purpose AI models owe documentation duties regardless of application [5], and a subset classified as posing systemic risk owe more [6].
United States federal policy has attached to compute, then withdrawn, then attached to use. Executive Order 14110 required reporting from developers of dual-use foundation models trained above
US state law has attached to a developer–deployer pair. Colorado’s SB 24-205 imposed duties on developers of high-risk AI systems — disclosures to deployers, documentation supporting impact assessments, notification of known discrimination risks — and separate duties on deployers, including risk-management programmes, impact assessments, consumer notice and appeal rights [13]. That statute was repealed and reenacted in May 2026 by SB 26-189, which recasts the regime around automated decision-making technology used in consequential decisions across education, employment, housing, financial services, insurance, health care and government benefits, with primary implementation from 1 January 2027 [14].
The United Kingdom attached to principles administered by existing regulators. Its white paper declined to create an AI-specific regulator, on the reasoning that a new cross-sector body would introduce complexity and undermine existing expert oversight, and set out five cross-sectoral principles — safety and robustness, appropriate transparency and explainability, fairness, accountability and governance, and contestability and redress — to be applied initially on a non-statutory basis [15].
China attaches to the service offered to the public. The 2023 Interim Measures cover the provision of text, image, audio and video generation services to the public in mainland China, expressly excluding internal research and development, and require services with public-opinion attributes or capacity for social mobilisation to undergo a security assessment and file their algorithm within ten business days of launch [16].
Set side by side, the consequences of each choice become legible. Attaching to a role produces clean allocation but depends on a self-declared intended purpose, so the classification rests on a statement rather than an observation. Attaching to a use case tracks harm most directly but requires a contestable judgement about the causal weight of an output — M-25-21 handles this by presuming certain categories high-impact and demanding written documentation to rebut the presumption [12]. Attaching to the service is the most robust to model churn, because a service can be inspected while it runs, but it reaches only what is offered publicly and misses internal or embedded deployments. Attaching to the model reaches the point of greatest leverage and least observability at once. And attaching to compute is administratively the easiest of all, which is why it keeps recurring, and why it deserves its own section.
Compute thresholds: administrable, and a poor proxy
The case for compute as a governance handle is not about capability at all. It is about enforceability. Compute is detectable, excludable and quantifiable, and it is produced through an extremely concentrated supply chain — properties that data and algorithms conspicuously lack [18]. A regulator can count chips. It cannot count insights.
Both major thresholds in circulation were built on that logic. The AI Act presumes a general-purpose AI model to have high-impact capabilities, and therefore systemic risk, when the cumulative training computation exceeds
The proxy is a weak one in three separate ways, and it is worth separating them.
First, the relationship between training compute and capability is not stable across architectures, data or training regimes. Hooker’s analysis of threshold-based governance argues directly that compute thresholds as currently implemented are shortsighted and likely to fail to mitigate risk, because the relationship between computational capacity and risk remains uncertain and fast-moving [17]. That is a contested position rather than a settled one; the compute-governance literature accepts the same uncertainty but treats compute as a usable lever precisely because nothing else is countable [18]. The disagreement is genuine and not resolvable from published evidence: it is a dispute about whether a bad proxy that can be enforced beats a good criterion that cannot.
Second, the meaning of a fixed number drifts, and quickly. Ho and colleagues, analysing more than two hundred language-model evaluations from 2012 to 2023, estimated that the compute required to reach a set performance threshold has halved roughly every eight months, with a 95% confidence interval of about five to fourteen months [19]. If a threshold is fixed at
The analysis that follows is mine rather than the authors’. Taking the central estimate at face value, a threshold left unamended for two years admits models with roughly the effective capability that eight times the compute bought when the number was written; at the pessimistic end of the interval it is far more. The drift runs in both directions at once. Capable systems slip under the threshold as efficiency improves, while ordinary systems cross over it as hardware gets cheaper and larger training runs become routine. Under-inclusion and over-inclusion grow together, which is why the delegated-act power in Article 51 is not an afterthought but the load-bearing part of the design [4].
Third, the quantity itself is ambiguous. The AI Act refers to the cumulative computation used for training [4], but a modern system’s capability is assembled from stages with very different compute signatures: pretraining, post-training, and distillation from a larger teacher. A distilled model can inherit much of the behaviour of a run it did not pay for. Whether the teacher’s compute counts toward the student’s total is a question of measurement convention, not of physics, and the convention determines who is regulated. This is analysis, not a claim that any regulator has resolved it incorrectly.
Documentation: what a file can evidence
The documentation instruments now written into law began as voluntary research proposals. Model cards proposed that released models be accompanied by short documents recording intended use, evaluation procedures, and disaggregated performance across groups [20]. Datasheets for datasets proposed the analogous artefact one level down, recording a dataset’s motivation, composition, collection process and recommended uses [21]. Both were arguments for a professional norm. Both have since been absorbed into regulatory text.
The AI Act’s general-purpose model regime is the clearest example. Providers must draw up and keep up to date technical documentation of the model including its training and testing process, make it available to the AI Office and national authorities on request, supply information and documentation to downstream providers integrating the model, and publish a sufficiently detailed summary of the content used for training according to a Commission template; models released under free and open-source licences with public parameters are largely exempt unless they carry systemic risk [5]. The accompanying General-Purpose AI Code of Practice, published on 10 July 2025 and structured in transparency, copyright, and safety-and-security chapters, supplies a model documentation form and functions as a voluntary route to demonstrating compliance with reduced administrative burden [8].
What such a file can genuinely evidence is narrower than it first appears. It can establish that a process was followed, that specific tests were run, that particular risks were considered, and that a named party asserted particular facts on a particular date — which is exactly what a liability regime needs, and no small thing. What it cannot establish is the behaviour of the system a user meets today. Documentation is a record of an instant, and the instant it records is the one at which the file was signed.
The empirical picture of what is actually disclosed is not encouraging, and it has moved in the wrong direction on at least one measurement. The 2025 Foundation Model Transparency Index, scoring thirteen major developers, reports that the average score out of 100 fell from 58 in 2024 to 40 in 2025, that companies are most opaque about training data and training compute and about post-deployment usage and impact, and that while companies do tend to disclose evaluations of capabilities and risks, those disclosures suffer from limited methodological transparency, limited third-party involvement, limited reproducibility, and limited reporting of train–test overlap [27]. The same index notes that signatories to the EU code score higher on average than non-signatories, alongside open-model and enterprise-focused developers [27]. That is a correlation across a small sample with obvious selection effects; it is not evidence that signing caused the disclosure, and the authors present it as a grouping rather than a causal finding.
Conformity assessment, and the audit that does not yet exist
Product-safety law conventionally checks compliance before market entry. The AI Act follows that template, and the detail of how it does so is frequently misread. For most high-risk systems the procedure is conformity assessment based on internal control — self-assessment, with no notified body involved. Third-party assessment is triggered in narrower circumstances, chiefly where harmonised standards do not exist, have not been applied, or have been applied only in part [3]. For systems intended for law enforcement, immigration or asylum authorities, the market surveillance authority acts as the notified body rather than a private one [3].
The general-purpose model layer is looser still. Article 55 requires providers of models with systemic risk to perform model evaluation in accordance with standardised protocols and tools, including adversarial testing, to assess and mitigate systemic risks at Union level, to document and report serious incidents to the AI Office without undue delay, and to maintain adequate cybersecurity for the model and its physical infrastructure [6]. Until harmonised standards are published, compliance may be demonstrated through an approved code of practice or by other means shown adequate to the Commission [6].
The phrase “standardised protocols and tools” is doing an enormous amount of work, and it is the honest place to say plainly that no settled audit methodology for these systems exists. The research literature is a set of competing proposals rather than a practice. Raji and colleagues set out an internal auditing framework running across the development lifecycle and producing documentation at each stage [22]. Mokander and colleagues propose a three-layered structure — governance audits of the provider, model audits after pretraining, and application audits — while acknowledging its limitations [23]. Casper and colleagues argue that black-box access is insufficient for rigorous audits, that white-box access to weights, activations and gradients and outside-the-box access to training and deployment information permit substantially more scrutiny, and that the level of access an auditor had must itself be disclosed for findings to be interpretable [24]. Weidinger and colleagues survey current safety evaluation and find capability evaluations to be the main approach in use, with identified gaps at the layers of human interaction and systemic impact, and note that context determines whether a capability causes harm at all [25]. Anderljung and colleagues propose registration and reporting requirements and external scrutiny as building blocks while arguing that self-regulation alone is insufficient [26].
These are five constructive proposals that do not agree on scope, on required access, or on what a passing result would mean. Compare the machinery a financial audit rests on — promulgated standards, licensed practitioners, defined materiality, sampling theory, statutory liability — and none of the equivalents exist here. Analysis: a conformity regime whose default is self-assessment against standards still being drafted is not weak regulation so much as a regime whose binding force is deferred until the standards land. The date on which harmonised standards are published matters more to what the AI Act requires in practice than the date on which its articles applied.
The general-purpose problem: duties that depend on facts you cannot see
The sharpest structural difficulty in the current designs is that obligations frequently land on a party who cannot observe the facts the obligations turn on.
A deployer building on a general-purpose model owes duties that depend on training data provenance, evaluation coverage, known failure modes and the presence or absence of systemic risk. Every one of those facts lives upstream. The AI Act’s answer is contractual and informational: Article 25 requires third parties supplying components, tools or services integrated into high-risk systems to agree in writing what information and technical access they will provide, with an exception for free and open-source components, and empowers the AI Office to develop voluntary model contract terms; where a deployer puts its trademark on a system, substantially modifies it, or changes its intended purpose such that it becomes high-risk, the deployer becomes the provider and the original provider must cooperate and supply necessary information [2]. Article 53 supplies the parallel duty from the model side, requiring general-purpose model providers to make information available to downstream integrators [5], and Article 3 names the downstream provider as a distinct role in the chain [1]. Colorado’s developer–deployer split has the same shape: developer disclosures are the mechanism by which a deployer becomes able to complete its own impact assessment [13].
Contract is a transmission mechanism. It is not a discovery mechanism. If the upstream party does not know something, or characterises it in terms the downstream party cannot evaluate, no clause makes the fact appear. The transparency index measurement bears directly on this: training data and training compute are the categories developers are most opaque about, and disclosed evaluations come with limited reproducibility and limited reporting of train–test overlap [27]. A deployer told that a model was evaluated for a risk cannot generally tell whether the evaluation would have detected it — which is precisely the gap Casper and colleagues identify when they argue that auditor access level determines what a finding can support [24].
The enforcement gap: a stable name over a moving system
The last problem is temporal, and it is the one for which no framework has a convincing answer.
The clearest published demonstration that a served system moves under a fixed name is Chen, Zaharia and Zou’s longitudinal comparison of GPT-3.5 and GPT-4 between March and June 2023, which found substantial behavioural shifts across mathematics, sensitive-question handling, code generation and visual reasoning — reporting, for example, that GPT-4 identified prime versus composite numbers at 84% accuracy in March and 51% in June [28]. The paper was contested. Narayanan and Kapoor argued that the results demonstrate behavioural change rather than capability decline, drawing the distinction that capabilities are acquired in expensive pretraining while behaviour is shaped by fine-tuning that happens repeatedly; on the primality task specifically they showed that the March and June models were equally poor once composite numbers were included in the test set, differing in calibration bias rather than in mathematical ability [29].
Both readings are defensible and this article takes neither as settled. What matters here is that both sides agree on the fact a regulator has to work with: the observable behaviour of a service accessed under one stable name differed materially between two dates, for reasons not disclosed at the time. Whether that constitutes a capability change or a behaviour change is a scientific question. For a rule that binds outputs, it is a distinction without a difference.
Set that against the cadence of the instruments themselves. The AI Act’s obligations phase in across years: prohibitions and AI literacy from February 2025, governance and general-purpose model rules from August 2025, the core rules from August 2026, high-risk systems in sensitive areas from December 2027, and embedded high-risk systems from August 2028 [7]. Colorado’s regime was postponed, then repealed and reenacted with implementation from 2027 [13, 14]. Executive Order 14110 was in force for under fifteen months [9, 10]. Meanwhile the artefact updates on a schedule measured in weeks.
No pre-market check can close that gap, and the frameworks partially know it. The provisions likely to bind hardest are the continuous ones rather than the point-in-time ones: incident tracking and reporting to the AI Office without undue delay [6], ongoing monitoring of high-impact use with an obligation to discontinue non-compliant AI [12], and the measure and manage functions of a risk framework designed to be applied throughout a lifecycle rather than at a gate [11].
Predictions, with the observations that would falsify them
These are forecasts, separated from the sourced analysis above. Horizon: 8 August 2029. Assumptions common to all four: no capability discontinuity that triggers emergency legislation; continued rapid model release cadence; the AI Act remains in force in substantially its current architecture.
One. Compute thresholds will be supplemented by capability-based triggers rather than abandoned, because the administrability argument survives every criticism of the proxy [18, 17]. Indicator: a delegated act or successor instrument that adds evaluation-based criteria alongside the floating-point figure. Disconfirmed if the numerical threshold is either removed outright with nothing put in its place, or left unamended with no evaluation-based criterion added.
Two. The operative centre of gravity will shift from pre-market documentation toward post-market monitoring and incident reporting. Indicator: the first substantial enforcement actions turn on failures to report or monitor rather than on defective technical files. Disconfirmed if the majority of publicly reported enforcement concerns documentation completeness at the point of market entry.
Three. Version identity will become a regulated object: obligations will increasingly attach to a specified served configuration, with change logs and re-assessment triggers, rather than to a product name. Indicator: standards or guidance that define what constitutes a substantial modification of a served model. Disconfirmed if guidance published by 2029 continues to treat a named model as a single regulated entity across its serving history.
Four. Third-party AI audit will remain non-standardised: no licensure regime with defined scope, materiality and liability will be operating at scale. Indicator: whether an accreditation scheme with mandatory access levels exists and has accredited practitioners. Disconfirmed if such a scheme is operating and required for any broad class of systems [24, 23].
What to take away
Inspection of weights and measures is the oldest working example of regulating an instrument you did not build, and it works because two things are true. There is a standard held somewhere other than the maker’s premises, against which any instrument can be compared. And there is a seal, applied to what has been checked, which breaks if the instrument is opened afterwards.
Current AI frameworks have made real progress on the first. Documentation duties, transparency indices and codes of practice are all attempts to construct something to compare against [5, 27, 8]. The second is missing almost entirely. A served system can be altered after assessment with no external mark that anything changed, and the review cycles of every instrument now in force are slower than the systems they name.
The practical consequence for anyone reading a compliance claim is a single question, and it is the same question in every jurisdiction: what exactly was checked, by whom, with what access, and on what date — and is the thing running now the thing that was checked. Where that question cannot be answered, what is in front of you is a record of an inspection rather than evidence about a system.