A regime is machinery, not a sentence
A companion piece in this series asked what an AI rule can actually bind — a model as artefact, a provider or deployer role, a use case, a training-compute threshold, or a service as offered — and found that each naming choice buys enforceability somewhere and loses it somewhere else. That question assumed a rule already exists and is trying to attach itself to an object. This article assumes the naming problem is solved and asks the next one, which is duller to state and more consequential in practice: once a rule has named its object, what actually happens to it?
The honest answer is that almost nothing happens directly. A statute that says a high-risk system “shall be designed and developed in such a way as to achieve an appropriate level of accuracy, robustness and cybersecurity” does not, by itself, tell anyone what accuracy number passes. It has to pass through a chain of institutional machinery before it touches an engineer’s desk: a classification step decides which of a small number of tiers applies; a standards-setting step turns the statute’s adjectives into numbers and procedures; a conformity-assessment step is where a system is checked — usually by its own provider — against those numbers; a market-surveillance and registration step keeps that determination visible after the product ships; and an incident-reporting and penalty step gives the whole apparatus consequences when something goes wrong. None of these five stages is optional, and none of them is what people usually mean when they say “the AI Act” or “AI regulation.” The regulation is the assembly of all five running together.
This article walks that machine end to end, using the EU AI Act as the running specimen because it is currently the most completely built example of all five stages operating at once, with the US National Institute of Standards and Technology’s voluntary framework and the international ISO/IEC 42001 certification standard brought in wherever they illuminate a different way of doing the same job outside a statute. It does not re-argue what a rule can bind or what a model card can evidence — that is the companion piece’s territory. It stays on the operating machinery: how a system gets sorted, how a requirement becomes checkable, how the check actually happens, how the system stays visible afterward, and how the whole thing bites.
Classification: what actually triggers each tier
The EU AI Act sorts systems into four tiers, and only two of them are triggered by an examination of the technology itself. The first tier, unacceptable risk, is a fixed list of prohibited practices rather than a threshold: manipulative techniques that materially distort behaviour and cause significant harm, exploitation of the vulnerabilities of specific groups, social scoring, biometric categorisation inferring protected characteristics, untargeted scraping of facial images to build recognition databases, emotion inference in workplaces and schools, and real-time remote biometric identification in public spaces by law enforcement outside narrow judicially authorised exceptions [1]. A ninth prohibition, on certain uses inferring emotion or targeting through subliminal manipulation in additional contexts, phases in later than the original eight [14]. There is no proportionality test inside this tier: a system that matches one of these descriptions is banned regardless of how well it works.
The second tier, high-risk, is where the machinery gets interesting, because it has two independent triggers. A system is high-risk if it is a safety component of a product already subject to third-party conformity assessment under existing EU product-safety law — toys, machinery, medical devices, lifts — or if it falls into one of eight domains listed in Annex III: biometrics, critical infrastructure, education and training, employment, access to essential private and public services, law enforcement, migration and border control, and the administration of justice and democratic processes [2, 3]. The second pathway carries a genuinely load-bearing escape hatch. A provider whose system falls within an Annex III category may still treat it as not high-risk if it performs a narrow procedural task, improves the output of a completed human activity, detects decision-making patterns without replacing human review, or performs a preparatory task — unless the system profiles natural persons, in which case the exemption never applies regardless of function [2]. Mechanically, this exemption is claimed the same way a great deal of this regime works: the provider documents its own reasoning before placing the system on the market and produces that documentation to a national authority only if asked. The classification of an entire category of systems can therefore rest on an assessment nobody outside the provider has reviewed at the point the system ships.
The third tier, transparency obligations, attaches not to a domain but to an interaction. Providers of systems that talk to people must make that fact detectable unless it is obvious to a reasonably well-informed person; systems generating synthetic audio, image, video or text must mark the output as machine-generated in a machine-readable way; and deployers exposing people to emotion-recognition or biometric-categorisation systems, or publishing AI-generated text on matters of public interest without human editorial review, carry a parallel disclosure duty [5]. Everything left over — the overwhelming majority of deployed systems — falls into the fourth, minimal-risk tier and carries no tier-specific obligation at all.
What makes classification a piece of machinery rather than a one-time definition is that the high-risk list is not fixed. Article 7 gives the Commission power to add to or remove from Annex III by delegated act, and it specifies exactly what has to be shown to do so: the candidate system must fall within an area Annex III already covers, and it must pose a risk to health, safety or fundamental rights at least equivalent to the risks already listed there, judged against ten named factors — the system’s intended purpose and reach, the nature and volume of data it processes, its level of autonomy and the possibility of human override, any documented history of harm, the severity and reversibility of possible harm, the vulnerability or power imbalance of those affected, and the adequacy of existing legal safeguards [4]. A category can also be removed if it no longer meets that bar without lowering the Act’s overall protection. The practical effect is that Annex III functions less like a fixed catalogue and more like a living index that the Commission can extend without going back to the legislature — which is precisely how a classification regime keeps pace with a technology that does not hold still, and precisely why what counts as high-risk today is not a permanent fact about a given use case.
From a legal duty to a checkable specification
A statute that requires “appropriate levels of accuracy, robustness and cybersecurity” has not yet told anyone what to build. Converting that language into something an engineer can check against is a distinct piece of machinery, and the EU AI Act performs it through Article 40: once the Commission issues a standardisation request, and a harmonised standard answering that request is published with its reference cited in the Official Journal, a high-risk system or general-purpose AI model that conforms to the standard is presumed to comply with the corresponding legal requirement — to the extent the standard actually covers it [6]. That last clause matters mechanically: coverage can be partial, so a single harmonised standard might discharge the presumption for part of a chapter’s requirements and leave the rest to be demonstrated another way. The Commission’s request itself has to specify technical compliance obligations, documentation and reporting processes, and — a detail worth noting on its own — resource-efficiency criteria including energy consumption, so the standard-setting brief is not confined to safety in the narrow sense [6].
The body actually answering that request is CEN-CENELEC’s Joint Technical Committee 21, established in 2021 and drawing more than three hundred experts from over twenty countries across five working groups [18]. Its output is not one document but a small family: a foundational AI trustworthiness framework, an AI risk-management standard, an AI quality-management-system standard, and an AI conformity-assessment standard, alongside more specific deliverables on datasets, bias, robustness and logging [18]. This is the concrete machinery behind the word “standard” as it appears in the statute: a named committee, working from a named request, publishing named documents that a manufacturer can cite by number.
It is worth setting this specific legal mechanism against two things that look similar and are not. The first is the US National Institute of Standards and Technology’s AI Risk Management Framework, organised around four functions — govern, map, measure and manage — that an organisation applies across an AI system’s lifecycle [15]. The framework is voluntary and carries no statutory presumption of compliance with anything; its role is to give organisations a common process vocabulary rather than a pass/fail bar tied to a specific law. NIST extends that vocabulary the same way a legislature might amend an annex, but through a different instrument: a companion “profile” document rather than a rule change. Its Generative AI Profile, published in July 2024, applies the four functions specifically to generative systems and organises the exercise around twelve named risk categories, including chemical, biological, radiological and nuclear information, confabulation, data privacy, intellectual property, information integrity and value-chain integration [16]. The mechanism on display is a general framework becoming actionable for one technology class through an added companion document, not through new legislation.
The second look-alike is ISO/IEC 42001, an international standard published in 2023 that specifies requirements for an organisation-wide AI management system, structurally analogous to the long-established ISO 9001 quality-management and ISO 27001 information-security standards [17]. Where a CEN-CENELEC harmonised standard buys a legal presumption of conformity inside a specific statute, ISO/IEC 42001 buys something else: a certificate, issued by an accredited third-party certification body after an audit, that functions as a market signal independent of any particular jurisdiction’s law. Microsoft’s own compliance documentation is a useful illustration of what that mechanism produces in practice, listing several of its AI-facing products — including GitHub Copilot and Microsoft 365 Copilot — as having undergone independent third-party audits against the standard, with certificates and audit reports published for customers to inspect [17]. Two systems can therefore both be described as “meeting a recognised AI standard” while running on entirely different levers: one anchored to a statute’s presumption-of-conformity clause, the other anchored to reputation and procurement preference, with no automatic legal effect under the AI Act as currently written.
Conformity assessment: where the check actually happens
Once a system is classified and a standard exists to check it against, something has to perform the check. Article 43 sets out two procedural pathways, and the split between them is the single most consequential fact about how EU AI governance actually operates day to day. For most Annex III categories — everything except biometrics — the default procedure is internal control under Annex VI: no notified body is involved at all. Third-party assessment under Annex VII is triggered only in narrower circumstances, chiefly where no harmonised standard exists for a requirement, where a provider has not applied an existing standard, or where a standard has been published subject to restriction [8].
It is worth being precise about what internal control actually requires someone to do, because the phrase can sound more rigorous than the mechanics turn out to be. Annex VI specifies three steps: confirm that the provider’s own quality management system complies with the separate obligations set out in Article 17; examine the technical documentation to assess whether the system meets the essential requirements in the relevant chapter; and verify that the design, development and post-market monitoring process is consistent with that same technical documentation [9]. Every one of those three checks is the provider comparing its own paperwork against its own paperwork. That is not a criticism so much as a literal description of what “self-assessment” means as a procedure — and it is why, for the majority of high-risk systems on the market, “this system has undergone conformity assessment” and “this system’s provider filled in a checklist about itself” describe the same event.
A notified body enters the picture only for the narrower set of cases where internal control is unavailable, and Article 31 specifies what an entity has to demonstrate before it can act as one: legal personality established under a member state’s law, documented independence from the providers and competitors it assesses, personnel who are not themselves involved in designing, marketing or consulting on the systems they evaluate, professional liability insurance unless the member state assumes that liability directly, and ongoing participation in EU-level coordination and standardisation activity [7]. A notified body, in other words, is a specific accredited institution operating under specific constraints — not a generic “auditor,” and not something a provider can simply hire on demand the way it can complete an internal checklist. One further wrinkle is worth naming because it collapses a distinction the rest of this article otherwise keeps: for systems intended for use by law-enforcement or migration and asylum authorities, the market-surveillance authority itself performs the notified-body role, meaning the same institution that will later police the system in the field also pre-certifies it before deployment [8].
None of this machinery was built from scratch for software. Veale and Zuiderveen Borgesius’s analysis of the Act situates its conformity-assessment structure squarely within the EU’s decades-old New Legislative Framework — the same self-declaration-plus-notified-body-plus-market-surveillance architecture used for toys, machinery and medical devices — and raises a concern, characterised rather than resolved, about whether an enforcement design built for inspecting physical products before they leave a factory transfers well onto software systems that keep changing after they ship, and about how much the whole edifice depends on national market-surveillance authorities whose capacity varies considerably across member states [20]. That is a scholarly concern about institutional capacity, not a finding that the architecture fails; it is included here because it explains why the machine has the shape it has — inherited, not purpose-built — rather than as a verdict on whether the inheritance was wise.
Market surveillance and the registry: the machine after the gate
Conformity assessment happens once, before a system reaches the market. Everything after that point runs through a different apparatus, and Article 74 is where its powers are specified. Market-surveillance authorities apply the EU’s general market-surveillance regulation to AI systems, can request access to training, validation and testing datasets through APIs or other technical means, can request source code as a last-resort measure once other means of verification are exhausted, and can conduct joint investigations across member states when a system presents risk that crosses borders [12]. Rather than creating one new AI-specific regulator, the Act mostly distributes this authority to bodies that already exist: financial supervisors retain oversight of AI used in financial services, data-protection authorities or other designated bodies oversee law-enforcement and justice systems, and the European Data Protection Supervisor covers AI used by the EU’s own institutions [12]. Enforcement capacity, in other words, is mostly reused rather than built new — which is the same institutional-capacity question Veale and Zuiderveen Borgesius raise about notified bodies, applied one layer further downstream.
The public-facing part of this ongoing apparatus is the EU database established under Article 71. Providers or their authorised representatives register the data specified in Annex VIII for systems classified as high-risk — and, notably, also for systems a provider has determined are not high-risk under the Article 6(3) exemption discussed above — while public-authority deployers register a separate section of data themselves [10]. Most of what is registered is publicly and freely browsable; the exception is data from real-world testing, which stays restricted to market-surveillance authorities and the Commission unless the provider consents to publish it, and the personal data the database holds is deliberately minimal, limited to the name and contact details of whoever is responsible for the registration entry [10]. This registry is the only part of the whole machine an outside party — a journalist, a researcher, an affected individual — can inspect without invoking a formal information-request power. Its completeness, however, depends entirely on providers correctly registering in the first place, including correctly registering their own claim that a system is exempt from high-risk treatment — which loops the registry’s reliability straight back to the same self-assessed classification step described earlier in this article.
Incident reporting and enforcement: the loop that closes
The two stages above are largely static: a classification made once, a registration filed once. What gives the machine a pulse after deployment is the duty to report serious incidents, and Article 73 specifies that duty with unusual precision. A provider must notify the market-surveillance authority of the member state where the incident occurred once it has established a causal link, or a reasonable likelihood of one, between its system and the incident — and the deadline this triggers depends on severity. The default window is fifteen days. Widespread infringement or serious and irreversible disruption to critical infrastructure compresses that to two days. A death compresses it further, to ten days, or immediately once a causal link is even suspected [11]. Providers may file an incomplete initial report followed by a complete one, must investigate the incident, assess the risk, take corrective action and cooperate with the relevant authorities and notified bodies, and the receiving authority itself then has seven days to take appropriate measures, with an obligation to notify the Commission immediately for the most serious cases [11]. This is the one part of the whole apparatus that runs continuously rather than at a single gate: a system that passed its conformity assessment cleanly can still trigger this clock on any day of its operating life.
What gives all of the preceding stages consequences is Article 99’s three-tier penalty structure. Breaching the prohibited-practices list carries fines up to the higher of thirty-five million euros or seven percent of global annual turnover; breaching most other provider, deployer or notified-body obligations caps at fifteen million euros or three percent; and supplying false, incomplete or misleading information to authorities caps at seven and a half million euros or one percent — with small and medium enterprises assessed against the lower percentage figures throughout, and authorities required to weigh severity, the operator’s size, cooperation, intent and any mitigating action taken before setting a fine [13]. Member states must report enforcement activity to the Commission annually and provide judicial remedies for those affected [13]. It is worth being clear about how a case actually reaches this stage in practice: penalties are the ceiling of a process that almost always starts somewhere further up this article — a self-reported incident, a market-surveillance investigation, a registry inconsistency, or a complaint — rather than something a regulator imposes by independently discovering a violation from nothing.
The soft layer underneath
One more piece of the machinery deserves a brief, deliberately limited mention, because comparing how different jurisdictions coordinate their AI regimes is a separate piece of this series and not the concern of this one. The OECD’s AI Principles, first adopted in 2019 and revised in 2024, are a non-binding Council Recommendation rather than a treaty, currently adhered to by forty-seven governments and institutions, including the European Union itself, organised around values — inclusive growth, human rights, transparency, robustness, accountability — and complementary recommendations for policymakers on research investment, ecosystem-building and international cooperation [19]. Their mechanical role in the picture this article has traced is modest but real: they function as shared vocabulary that later instruments frequently echo — NIST’s four RMF functions and the EU Act’s own list of essential requirements both draw on language this document popularised — without any of those later instruments being legally derived from it. It is the substrate the harder machinery sits on, not a stage in the machinery itself.
Predictions, with the observations that would falsify them
These are forecasts, separated from the sourced description above. Horizon: 16 August 2029. Assumptions common to all four: the EU AI Act remains in force in substantially its current architecture; no capability discontinuity triggers emergency legislation; the Commission continues issuing standardisation requests on broadly its current cadence.
One. Internal control under Annex VI will remain the pathway for the large majority of registered high-risk systems, not because harmonised standards fail to appear but because most Annex III categories are legally eligible for it once an applicable standard exists. Indicator: published Commission or AI Office statistics comparing notified-body assessments to self-declared internal-control registrations in the EU database. Disconfirmed if notified-body assessments come to account for a majority of registered high-risk systems.
Two. CEN-CENELEC’s foundational standards — the trustworthiness framework, risk-management, quality-management and conformity-assessment documents — will be formally cited in the Official Journal, and therefore trigger the presumption of conformity, later than the Commission’s original standardisation-request schedule anticipated. Indicator: the date of first Official Journal citation of a JTC 21 AI Act harmonised standard. Disconfirmed if the core standards are cited on or ahead of the Commission’s originally requested timeline.
Three. Voluntary ISO/IEC 42001 certification uptake will keep growing as a de facto procurement signal even though the AI Act attaches no direct legal effect to it. Indicator: growth in publicly listed ISO/IEC 42001 certificates among major model and system providers. Disconfirmed if enterprise procurement standardises instead on a different mechanism, such as a conformity mark created specifically under the AI Act, and ISO/IEC 42001 adoption plateaus or declines among frontier providers.
Four. Early enforcement activity will concentrate among pre-existing sectoral authorities — financial supervisors, data-protection authorities — rather than among newly designated general AI market-surveillance authorities, because the former already have investigatory infrastructure and case law while the latter must build both from nothing. Indicator: the distribution of publicly reported enforcement actions by authority type in the years after the core obligations phase in. Disconfirmed if newly designated general AI market-surveillance authorities generate the majority of early enforcement actions.
What to take away
None of the five stages traced here is, by itself, what a rule “is.” Classification decides what applies, and the list it draws on can be rewritten by delegated act without new legislation. Standards translate a statute’s adjectives into numbers, through at least three distinct institutional routes — a harmonised standard tied to a specific legal presumption, a voluntary framework with no such tie, and a certifiable management standard sold as a market signal — that are easy to mistake for one another and are not. Conformity assessment is where a system meets those numbers, and for most systems that meeting is a provider checking its own file against its own checklist, with a notified body appearing only at the edges. The registry and market-surveillance powers keep the result visible, imperfectly, after the gate. And incident reporting and penalties are what make the whole arrangement matter rather than describe.
The practical habit this suggests is the same one the companion piece on binding arrived at from a different angle: before treating a compliance claim as settled, ask which of these five stages actually produced it, who performed the check, and whether that stage runs once at a gate or continuously across the system’s life. A great deal of what gets called “AI regulation” is still, as of this writing, machinery mid-assembly — standards not yet cited, registrations not yet due, penalty regimes not yet tested by a real case. Knowing which gears are turning and which are still being bolted into place is the difference between reading a finished mechanism and reading a claim about one.